Improving word embedding factorization for compression using distilled nonlinear neural decomposition

10Citations
Citations of this article
70Readers
Mendeley users who have this article in their library.

Abstract

Word-embeddings are vital components of Natural Language Processing (NLP) models and have been extensively explored. However, they consume a lot of memory which poses a challenge for edge deployment. Embedding matrices, typically, contain most of the parameters for language models and about a third for machine translation systems. In this paper, we propose Distilled Embedding, an (input/output) embedding compression method based on low-rank matrix decomposition and knowledge distillation. First, we initialize the weights of our decomposed matrices by learning to reconstruct the full pre-trained word-embedding and then fine-tune end-to-end, employing knowledge distillation on the factorized embedding. We conduct extensive experiments with various compression rates on machine translation and language modeling, using different data-sets with a shared word-embedding matrix for both embedding and vocabulary projection matrices. We show that the proposed technique is simple to replicate, with one fixed parameter controlling compression size, has higher BLEU score on translation and lower perplexity on language modeling compared to complex, difficult to tune state-of-the-art methods.

Cite

CITATION STYLE

APA

Lioutas, V., Rashid, A., Kumar, K., Akmal Haidar, M., & Rezagholizadeh, M. (2020). Improving word embedding factorization for compression using distilled nonlinear neural decomposition. In Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020 (pp. 2774–2784). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.findings-emnlp.250

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free