Abstract
Hallucination occurs when large language mod-els exhibit behavior that deviates from the boundaries of their knowledge during response generation. To address this critical issue, pre-vious learning-based methods attempt to fine-tune models but are limited by off-policy sam-pling and coarse-grained feedback. In this paper, we present Reinforcement Learning for H allucination (RLFH), an on-policy self-alignment approach that enables LLMs to ac-tively explore their knowledge boundaries and self-correct generation behavior through fine-grained feedback signals. RLFH introduces a self-assessment framework where the policy serves as its own judge. Through this frame-work, responses are automatically decomposed into atomic facts and their truthfulness and informativeness are assessed against external knowledge sources. The resulting fine-grained feedback at the statement level are then con-verted into token-level dense reward signals. This enables online reinforcement learning to achieve precise and timely optimization with-out human intervention. Comprehensive eval-uations on HotpotQA, SQuADv2, and Biogra-phy benchmarks validate RLFH's effectiveness in hallucination mitigation.
Cite
CITATION STYLE
Wen, X., Lou, J., Lu, X., Yuqiu, J., Guan, X., Lu, Y., … Sun, L. (2025). On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 5215–5231). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.271
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.