Abstract
Modern Language Models (LMs) are capable of following long and complex instructions that enable a large and diverse set of user requests. While Information Retrieval (IR) models use these LMs as the backbone of their architectures, virtually none of them allow users to provide detailed instructions alongside queries, thus limiting their ability to satisfy complex information needs. In this work, we study the use of instructions in IR systems. We build FOLLOWIR, a rigorous instruction evaluation benchmark for following real-world instructions in IR. FOLLOWIR repurposes detailed instructions-also known as narratives-developed for professional assessors to evaluate retrieval systems. In particular, we build our benchmark from three collections curated for shared tasks at the Text REtrieval Conference (TREC). Through this process, we can measure how well IR models follow instructions, through a new pairwise evaluation framework. Our results indicate that existing retrieval models fail to correctly use instructions, using them for basic keywords and struggling to understand long-form information. However, we show that it is possible for IR models to learn to follow complex instructions: our new FOLLOWIR-7B model has significant improvements after fine-tuning on our training set.
Cite
CITATION STYLE
Weller, O., Chang, B., MacAvaney, S., Lo, K., Cohan, A., Van Durme, B., … Soldaini, L. (2025). FOLLOWIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions. In Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025 (Vol. 1, pp. 11926–11942). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.naacl-long.597
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.