Quantifying the effects of text duplication on semantic models

25Citations
Citations of this article
98Readers
Mendeley users who have this article in their library.

Abstract

Duplicate documents are a pervasive problem in text datasets and can have a strong effect on unsupervised models. Methods to remove duplicate texts are typically heuristic or very expensive, so it is vital to know when and why they are needed. We measure the sensitivity of two latent semantic methods to the presence of different levels of document repetition. By artificially creating different forms of duplicate text we confirm several hypotheses about how repeated text impacts models. While a small amount of duplication is tolerable, substantial over-representation of subsets of the text may overwhelm meaningful topical patterns.

Cite

CITATION STYLE

APA

Schofield, A., Thompson, L., & Mimno, D. (2017). Quantifying the effects of text duplication on semantic models. In EMNLP 2017 - Conference on Empirical Methods in Natural Language Processing, Proceedings (pp. 2737–2747). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/d17-1290

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free