Cache-augmented latent topic language models for speech retrieval

0Citations
Citations of this article
69Readers
Mendeley users who have this article in their library.

Abstract

We aim to improve speech retrieval performance by augmenting traditional N-gram language models with different types of topic context. We present a latent topic model framework that treats documents as arising from an underlying topic sequence combined with a cache-based repetition model. We analyze our proposed model both for its ability to capture word repetition via the cache and for its suitability as a language model for speech recognition and retrieval. We show this model, augmented with the cache, captures intuitive repetition behavior across languages and exhibits lower perplexity than regular LDA on held out data in multiple languages. Lastly, we show that our joint model improves speech retrieval performance beyond N-grams or latent topics alone, when applied to a term detection task in all languages considered.

Cite

CITATION STYLE

APA

Wintrode, J. (2015). Cache-augmented latent topic language models for speech retrieval. In NAACL-HLT 2015 - 2015 Student Research Workshop (SRW) at the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings (pp. 1–8). Association for Computational Linguistics (ACL). https://doi.org/10.3115/v1/n15-2001

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free