Needle in a Haystack: Label-Efficient Evaluation under Extreme Class Imbalance

4Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Important tasks like record linkage and extreme classification demonstrate extreme class imbalance, with 1 minority instance to every 1 million or more majority instances. Obtaining a sufficient sample of all classes, even just to achieve statistically-significant evaluation, is so challenging that most current approaches yield poor estimates or incur impractical cost. Where importance sampling has been levied against this challenge, restrictive constraints are placed on performance metrics, estimates do not come with appropriate guarantees, or evaluations cannot adapt to incoming labels. This paper develops a framework for online evaluation based on adaptive importance sampling. Given a target performance metric and model for p(y|x), the framework adapts a distribution over items to label in order to maximize statistical precision. We establish strong consistency and a central limit theorem for the resulting performance estimates, and instantiate our framework with worked examples that leverage Dirichlet-tree models. Experiments demonstrate an average MSE superior to state-of-the-art on fixed label budgets.

References Powered by Scopus

XGBoost: A scalable tree boosting system

32588Citations
N/AReaders
Get full text

Multidimensional Binary Search Trees Used for Associative Searching

5516Citations
N/AReaders
Get full text

Simulation and the Monte Carlo Method: Third Edition

989Citations
N/AReaders
Get full text

Cited by Powered by Scopus

Fast Bayesian optimization of Needle-in-a-Haystack problems using zooming memory-based initialization (ZoMBI)

14Citations
N/AReaders
Get full text

Revisiting multi-dimensional classification from a dimension-wise perspective

2Citations
N/AReaders
Get full text

Model ChangeLists: Characterizing Updates to ML Models

0Citations
N/AReaders
Get full text

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Cite

CITATION STYLE

APA

Marchant, N. G., & Rubinstein, B. I. P. (2021). Needle in a Haystack: Label-Efficient Evaluation under Extreme Class Imbalance. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1180–1190). Association for Computing Machinery. https://doi.org/10.1145/3447548.3467435

Readers' Seniority

Tooltip

PhD / Post grad / Masters / Doc 7

100%

Readers' Discipline

Tooltip

Engineering 3

43%

Computer Science 2

29%

Physics and Astronomy 1

14%

Economics, Econometrics and Finance 1

14%

Save time finding and organizing research with Mendeley

Sign up for free