Constructing test collections using multi-armed bandits and active learning

14Citations
Citations of this article
13Readers
Mendeley users who have this article in their library.
Get full text

Abstract

While test collections provide the cornerstone of system-based evaluation in information retrieval, human relevance judging has become prohibitively expensive as collections have grown ever larger. Consequently, intelligently deciding which documents to judge has become increasingly important. We propose a two-phase approach to intelligent judging across topics which does not require document rankings from a shared task. In the first phase, we dynamically select the next topic to judge via a multi-armed bandit method. In the second phase, we employ active learning to select which document to judge next for that topic. Experiments on three TREC collections (varying scarcity of relevant documents) achieve τ ≈ 0.90 correlation for P@10 ranking and find 90% of the relevant documents at 48% of the original budget. To support reproducibility and follow-on work, we have shared our code online1.

Cite

CITATION STYLE

APA

Rahman, M. M., Kutlu, M., & Lease, M. (2019). Constructing test collections using multi-armed bandits and active learning. In The Web Conference 2019 - Proceedings of the World Wide Web Conference, WWW 2019 (pp. 3158–3164). Association for Computing Machinery, Inc. https://doi.org/10.1145/3308558.3313675

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free