An asymptotic model for the English hapax/vocabulary ratio

14Citations
Citations of this article
93Readers
Mendeley users who have this article in their library.

Abstract

In the known literature, hapax legomena in an English text or a collection of texts roughly account for about 50% of the vocabulary. This sort of constancy is baffling. The 100-millionword British National Corpus was used to study this phenomenon. The result reveals that the hapax/vocabulary ratio follows a U-shaped pattern. Initially, as the size of text increases, the hapax/vocabulary ratio decreases; however, after the text size reaches about 3,000,000 words, the hapax/vocabulary ratio starts to increase steadily. A computer simulation shows that as the text size continues to increase, the hapax/vocabulary ratio would approach 1. © 2010 Association for Computational Linguistics.

Cite

CITATION STYLE

APA

Fengxiang, F. (2010). An asymptotic model for the English hapax/vocabulary ratio. Computational Linguistics, 36(4), 631–637. https://doi.org/10.1162/coli_a_00013

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free