Binary classification with corrupted labels

6Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

In a binary classification problem where the goal is to fit an accurate predictor, the presence of corrupted labels in the training data set may create an additional challenge. However, in settings where likelihood maximization is poorly behaved—for example, if positive and negative labels are perfectly separable—then a small fraction of corrupted labels can improve performance by ensuring robustness. In this work, we establish that in such settings, corruption acts as a form of regularization, and we compute precise upper bounds on estimation error in the presence of corruptions. Our results suggest that the presence of corrupted data points is beneficial only up to a small fraction of the total sample, scaling with the square root of the sample size.

Author supplied keywords

Cite

CITATION STYLE

APA

Lee, Y., & Barber, R. F. (2022). Binary classification with corrupted labels. Electronic Journal of Statistics, 16, 1367–1392. https://doi.org/10.1214/22-EJS1987

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free