Abstract
One of the research fields significantly affected by the emergence of "big data" is computational linguistics. A prominent example of a large dataset targeting this domain is the collection of Google Books Ngrams, made freely available, for several languages, in July 2009. There are two problems with Google Books Ngrams; the textual format (compressed with Deflate) in which they are distributed is highly inefficient; we are not aware of any tool facilitating search over those data, apart from the Google viewer, which, as a Web tool, has seriously limited use. In this paper we present a simple preprocessing scheme for Google Books Ngrams, enabling also search for an arbitrary n-gram (i.e., its associated statistics) in average time below 0.2 ms. The obtained compression ratio, with Deflate (zip) left as the backend coder, is over 3 times higher than in the original distribution.
Author supplied keywords
Cite
CITATION STYLE
Grabowski, S., & Swacha, J. (2012). Google books ngrams recompressed and searchable. Foundations of Computing and Decision Sciences, 37(4), 271–281. https://doi.org/10.2478/v10209-011-0015-8
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.