Good-Turing Frequency Estimation Without Tears

218Citations
Citations of this article
103Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Linguists and speech researchers who use statistical methods often need to estimate the frequency of some type of item in a population containing items of various types. A common approach is to divide the number of cases observed in a sample by the size of the sample; sometimes small positive quantities are added to divisor and dividend in order to avoid zero estimates for types missing from the sample. These approaches are obvious and simple, but they lack principled justification, and yield estimates that can be wildly inaccurate. I.J. Good and Alan Turing developed a family of theoretically well-founded techniques appropriate to this domain. Some versions of the Good-Turing approach are very demanding computationally, but we define a version, the Simple Good-Turing estimator, which is straightforward to use. Tested on a variety of natural-language-related data sets, the Simple Good-Turing estimator performs well, absolutely and relative both to the approaches just discussed and to other, more sophisticated techniques. © 1995, Taylor & Francis Group, LLC. All rights reserved.

Cite

CITATION STYLE

APA

Gale, W. A., & Sampson, G. (1995). Good-Turing Frequency Estimation Without Tears. Journal of Quantitative Linguistics, 2(3), 217–237. https://doi.org/10.1080/09296179508590051

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free