Identifying Incorrect Labels in the CoNLL-2003 Corpus

27Citations
Citations of this article
64Readers
Mendeley users who have this article in their library.

Abstract

The CoNLL-2003 corpus for English-language named entity recognition (NER) is one of the most influential corpora for NER model research. A large number of publications, including many landmark works, have used this corpus as a source of ground truth for NER tasks. In this paper, we examine this corpus and identify over 1300 incorrect labels (out of 35089 in the corpus). In particular, the number of incorrect labels in the test fold is comparable to the number of errors that state-of-the-art models make when running inference over this corpus. We describe the process by which we identified these incorrect labels, using novel variants of techniques from semi-supervised learning. We also summarize the types of errors that we found, and we revisit several recent results in NER in light of the corrected data. Finally, we show experimentally that our corrections to the corpus have a positive impact on three state-ofthe-art models.

References Powered by Scopus

GloVe: Global vectors for word representation

27161Citations
N/AReaders
Get full text

Neural architectures for named entity recognition

2603Citations
N/AReaders
Get full text

An introduction to conditional random fields

772Citations
N/AReaders
Get full text

Cited by Powered by Scopus

Learning from Noisy Labels for Entity-Centric Information Extraction

37Citations
N/AReaders
Get full text

Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future

17Citations
N/AReaders
Get full text

Detecting Label Errors by using Pre-Trained Language Models

10Citations
N/AReaders
Get full text

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Cite

CITATION STYLE

APA

Reiss, F., Xu, H., Cutler, B., Muthuraman, K., & Eichenberger, Z. (2020). Identifying Incorrect Labels in the CoNLL-2003 Corpus. In CoNLL 2020 - 24th Conference on Computational Natural Language Learning, Proceedings of the Conference (pp. 215–226). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.conll-1.16

Readers' Seniority

Tooltip

PhD / Post grad / Masters / Doc 18

75%

Researcher 4

17%

Lecturer / Post doc 2

8%

Readers' Discipline

Tooltip

Computer Science 21

72%

Linguistics 5

17%

Social Sciences 2

7%

Philosophy 1

3%

Save time finding and organizing research with Mendeley

Sign up for free