LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information

2Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Recent advancements in large language models (LLMs) have markedly improved their capacity to handle long text inputs; however, current models, including GPT-4o, still exhibit unsatisfactory performance in long-form generation. Generating high-quality long-form content still remains a significant challenge. In this paper, we present LongDPO, a novel approach designed to enhance long-form text generation through step-level supervision. By leveraging Monte Carlo Tree Search (MCTS) to collect stepwise preference pairs and employing a global memory pool to maintain factual accuracy, LongDPO effectively mitigates issues such as inconsistencies that are prevalent in long-context LLMs. Furthermore, we integrate critique-augmented generation to refine the selected preference pairs. Following the collection of stepwise preference pairs, we apply stepwise preference learning for fine-grained optimization. Experimental results demonstrate that our method enhances performance on long-form generation benchmarks (e.g. LongBench-Write) while maintaining nearly lossless performance on several general benchmarks.

Cite

CITATION STYLE

APA

Ping, B., Zeng, J., Meng, F., Wang, S., Zhou, J., & Zhang, S. (2025). LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 7613–7632). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.395

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free