CA-GAR: Context-Aware Alignment of LLM Generation for Document Retrieval

2Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Information retrieval has evolved from traditional sparse and dense retrieval methods to approaches driven by large language models (LLMs). Recent techniques, such as Generation-Augmented Retrieval (GAR) and Generative Document Retrieval (GDR), leverage LLMs to enhance retrieval but face key challenges: GAR's generated content may not always align with the target document corpus, while GDR limits the generative capacity of LLMs by constraining outputs to predefined document identifiers. To address these issues, we propose Context-Aware Generation-Augmented Retrieval (CA-GAR), which enhances LLMs by integrating corpus information into their generation process. CA-GAR optimizes token selection by incorporating relevant document information and leverages a Distribution Alignment Strategy to extract corpus information using a lexicon-based approach. Experimental evaluations on seven tasks from the BEIR benchmark and four non-English languages from Mr.TyDi demonstrate that CA-GAR outperforms existing methods.

Cite

CITATION STYLE

APA

Yu, H., Kang, J., Li, R., Liu, Q., He, L., Huang, Z., … Lu, J. (2025). CA-GAR: Context-Aware Alignment of LLM Generation for Document Retrieval. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 5836–5849). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.303

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free