Orthographic errors in Web pages: Toward cleaner Web corpora

39Citations
Citations of this article
99Readers
Mendeley users who have this article in their library.

Abstract

Since the Web by far represents the largest public repository of natural language texts, recent experiments, methods, and tools in the area of corpus linguistics often use the Web as a corpus. For applications where high accuracy is crucial, the problem has to be faced that a non-negligible number of orthographic and grammatical errors occur in Web documents. In this article we investigate the distribution of orthographic errors of various types in Web pages. As a by-product, methods are developed for efficiently detecting erroneous pages and for marking orthographic errors in acceptable Web documents, reducing thus the number of errors in corpora and linguistic knowledge bases automatically retrieved from the Web. © 2006 Association for Computational Linguistics.

Cite

CITATION STYLE

APA

Ringlstetter, C., Schulz, K. U., & Mihov, S. (2006). Orthographic errors in Web pages: Toward cleaner Web corpora. Computational Linguistics, 32(3), 295–340. https://doi.org/10.1162/coli.2006.32.3.295

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free