Compact, efficient and unlimited capacity: Language modeling with compressed suffix trees

11Citations
Citations of this article
100Readers
Mendeley users who have this article in their library.

Abstract

Efficient methods for storing and querying language models are critical for scaling to large corpora and high Markov orders. In this paper we propose methods for modeling extremely large corpora without imposing a Markov condition. At its core, our approach uses a succinct index - a compressed suffix tree - which provides near optimal compression while supporting efficient search. We present algorithms for on-the-fly computation of probabilities under a Kneser-Ney language model. Our technique is exact and although slower than leading LM toolkits, it shows promising scaling properties, which we demonstrate through oo-order modeling over the full Wikipedia collection.

Cite

CITATION STYLE

APA

Shareghi, E., Petri, M., Haffari, G., & Conn, T. (2015). Compact, efficient and unlimited capacity: Language modeling with compressed suffix trees. In Conference Proceedings - EMNLP 2015: Conference on Empirical Methods in Natural Language Processing (pp. 2409–2418). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/d15-1288

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free