Active learning and negative evidence for language identification

4Citations
Citations of this article
49Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Language identification (LID), the task of determining the natural language of a given text, is an essential first step in most NLP pipelines. While generally a solved problem for documents of sufficient length and languages with ample training data, the proliferation of mi-croblogs and other social media has made it increasingly common to encounter use-cases that don’t satisfy these conditions. In these situations, the fundamental difficulty is the lack of, and cost of gathering, labeled data: unlike some annotation tasks, no single “expert” can quickly and reliably identify more than a handful of languages. This leads to a natural question: can we gain useful information when annotators are only able to rule out languages for a given document, rather than supply a positive label? What are the optimal choices for gathering and representing such negative evidence as a model is trained? In this paper, we demonstrate that using negative evidence can improve the performance of a simple neural LID model. This improvement is sensitive to policies of how the evidence is represented in the loss function, and for deciding which annotators to employ given the instance and model state. We consider simple policies and report experimental results that indicate the optimal choices for this task. We conclude with a discussion of future work to determine if and how the results generalize to other classification tasks.

Cite

CITATION STYLE

APA

Lippincott, T., & Van Durme, B. (2021). Active learning and negative evidence for language identification. In DaSH-LA 2021 - 2nd Workshop on Data Science with Human-in-the-Loop: Language Advances, Proceedings (pp. 47–51). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2021.dash-1.8

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free