Abstract
This study investigates the impact of including LLM-rewritten documents and LLM-expanded queries in training data on the performance and ranking behavior of retrieval models for ad-hoc retrieval tasks. Specifically, we train models using various training datasets that combine human-written documents with LLM-rewritten ones, and human-generated queries with LLM-expanded ones. Hereafter, we refer to the latter types as LLM-modified documents and LLM-modified queries. This allows us to examine: (1) the performance difference between models trained on LLM-modified documents and those trained on human-generated documents; (2) the tendency of models trained with LLM-modified documents to rank specific document types higher; (3) the retrieval performance on human-generated queries for models trained using LLM-modified queries; and (4) the ranking preference of models trained with LLM-modified queries. Experimental results show that training with LLM-modified documents generally yields retrieval performance comparable to models trained solely on human-generated documents. However, regarding ranking behavior, models trained on LLM-modified documents exhibit a clear tendency to rank LLM-modified documents higher within mixed corpora. When training with LLM-modified queries, performance on human-generated queries degrades, possibly due to the mismatch in query distributions. We found that this source bias is similarly introduced when training with LLM-modified queries.
Author supplied keywords
Cite
CITATION STYLE
Nakachi, Y., & Kato, M. P. (2025). Impact of LLM-Modified Queries and Documents in Training Data on Neural Retrieval Models. In SIGIR-AP 2025 - Proceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (pp. 364–373). Association for Computing Machinery, Inc. https://doi.org/10.1145/3767695.3769485
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.