Google books ngrams recompressed and searchable

2Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

One of the research fields significantly affected by the emergence of "big data" is computational linguistics. A prominent example of a large dataset targeting this domain is the collection of Google Books Ngrams, made freely available, for several languages, in July 2009. There are two problems with Google Books Ngrams; the textual format (compressed with Deflate) in which they are distributed is highly inefficient; we are not aware of any tool facilitating search over those data, apart from the Google viewer, which, as a Web tool, has seriously limited use. In this paper we present a simple preprocessing scheme for Google Books Ngrams, enabling also search for an arbitrary n-gram (i.e., its associated statistics) in average time below 0.2 ms. The obtained compression ratio, with Deflate (zip) left as the backend coder, is over 3 times higher than in the original distribution.

Cite

CITATION STYLE

APA

Grabowski, S., & Swacha, J. (2012). Google books ngrams recompressed and searchable. Foundations of Computing and Decision Sciences, 37(4), 271–281. https://doi.org/10.2478/v10209-011-0015-8

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free