Abstract
Large language models (LLMs) have demonstrated impressive capabilities in generating human-like text and have been shown to store factual knowledge within their extensive parameters. However, models like ChatGPT can still actively or passively generate false or misleading information, increasing the challenge of distinguishing between human-created and machine-generated content. This poses significant risks to the authenticity and reliability of digital communication. This work aims to enhance retrieval models’ ability to identify the authenticity of texts generated by large language models, with the goal of improving the truthfulness of retrieved texts and reducing the harm of false information in the era of large models. Our contributions include: (1) we construct a diverse dataset of authentic human-authored texts and highly deceptive AI-generated texts from various domains; (2) we propose a self-supervised training method, RetrieverGuard, that enables the model to capture textual rules and styles of false information from the corpus without human labelled data, achieving higher accuracy and robustness in identifying misleading and highly deceptive AI-generated content.
Cite
CITATION STYLE
Chen, C., & Zhang, S. (2025). RetrieverGuard: Empowering Information Retrieval to Combat LLM-Generated Misinformation. In 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025 (pp. 4399–4411). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-naacl.249
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.