Target-side word segmentation strategies for neural machine translation

53Citations
Citations of this article
104Readers
Mendeley users who have this article in their library.

Abstract

For efficiency considerations, state-of-the-art neural machine translation (NMT) requires the vocabulary to be restricted to a limited-size set of several thousand symbols. This is highly problematic when translating into inflected or compounding languages. A typical remedy is the use of subword units, where words are segmented into smaller components. Byte pair encoding, a purely corpus-based approach, has proved effective recently. In this paper, we investigate word segmentation strategies that incorporate more linguistic knowledge. We demonstrate that linguistically informed target word segmentation is better suited for NMT, leading to improved translation quality on the order of magnitude of +0.5 BLEU and -0.9 TER for a medium-scale English?German translation task. Our work is important in that it shows that linguistic knowledge can be used to improve NMT results over results based only on the language-agnostic byte pair encoding vocabulary reduction technique.

Cite

CITATION STYLE

APA

Huck, M., Riess, S., & Fraser, A. (2017). Target-side word segmentation strategies for neural machine translation. In WMT 2017 - 2nd Conference on Machine Translation, Proceedings (pp. 56–67). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w17-4706

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free