Abstract
In this paper we improve over the hierarchical Pitman-Yor processes language model in a cross-domain setting by adding skipgrams as features. We find that adding skipgram features reduces the perplexity. This reduction is substantial when models are trained on a generic corpus and tested on domain-specific corpora. We also find that within-domain testing and crossdomain testing require different backoff strategies. We observe a 30-40% reduction in perplexity in a cross-domain language modelling task, and up to 6% reduction in a within-domain experiment, for both English and Flemish-Dutch.
Cite
CITATION STYLE
Onrust, L., Van Den Bosch, A., & Van Hamme, H. (2016). Improving cross-domain n-gram language modelling with skipgrams. In 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016 - Short Papers (pp. 137–142). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p16-2023
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.