Improving cross-domain n-gram language modelling with skipgrams

3Citations
Citations of this article
96Readers
Mendeley users who have this article in their library.

Abstract

In this paper we improve over the hierarchical Pitman-Yor processes language model in a cross-domain setting by adding skipgrams as features. We find that adding skipgram features reduces the perplexity. This reduction is substantial when models are trained on a generic corpus and tested on domain-specific corpora. We also find that within-domain testing and crossdomain testing require different backoff strategies. We observe a 30-40% reduction in perplexity in a cross-domain language modelling task, and up to 6% reduction in a within-domain experiment, for both English and Flemish-Dutch.

Cite

CITATION STYLE

APA

Onrust, L., Van Den Bosch, A., & Van Hamme, H. (2016). Improving cross-domain n-gram language modelling with skipgrams. In 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016 - Short Papers (pp. 137–142). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p16-2023

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free