Training a Neural Network in a Low-Resource Setting on Automatically Annotated Noisy Data

Michael A. Hedderich; Dietrich Klakow

Conference ProceedingsOPEN ACCESS

Training a Neural Network in a Low-Resource Setting on Automatically Annotated Noisy Data

Proceedings of the Annual Meeting of the Association for Computational Linguistics (2018) 12-18

DOI: 10.18653/v1/w18-3402

22Citations

120Readers

Abstract

Manually labeled corpora are expensive to create and often not available for low-resource languages or domains. Automatic labeling approaches are an alternative way to obtain labeled data in a quicker and cheaper way. However, these labels often contain more errors which can deteriorate a classifier's performance when trained on this data. We propose a noise layer that is added to a neural network architecture. This allows modeling the noise and train on a combination of clean and noisy data. We show that in a low-resource NER task we can improve performance by up to 35% by using additional, noisy data and handling the noise.

References Powered by Scopus

View more at Scopus

Cited by Powered by Scopus

View more at Scopus

Cite

CITATION STYLE

APA

Hedderich, M. A., & Klakow, D. (2018). Training a Neural Network in a Low-Resource Setting on Automatically Annotated Noisy Data. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 12–18). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-3402

Readers over time

Readers' Seniority

PhD / Post grad / Masters / Doc 36

69%

Researcher 12

23%

Lecturer / Post doc 3

Professor / Associate Prof. 1

Readers' Discipline

Computer Science 52

81%

Engineering 5

Linguistics 5

Business, Management and Accounting 2

Training a Neural Network in a Low-Resource Setting on Automatically Annotated Noisy Data

Abstract

References Powered by Scopus

Long Short-Term Memory

GloVe: Global vectors for word representation

Learning from noisy large-scale datasets with minimal supervision

Cited by Powered by Scopus

OCR on-the-go: Robust end-to-end systems for reading license plates & street signs

Analysing the Noise Model Error for Realistic Noisy Label Data

Handling noisy labels for robustly learning from self-training data for low-resource sequence labeling

Register to see more suggestions

Cite

Readers over time

Readers' Seniority

Readers' Discipline