Abstract
With growing applications of Machine Learning in daily lives, Natural Language Processing (NLP) has emerged as a heavily researched area. Finding its applications in tasks ranging from simple Q/A chatbots to fully fledged conversational AI, NLP models are vital. Word and Sentence embeddings are one of the most common starting points of any NLP task. A word embedding represents a given word in a predefined vector-space while maintaining vector relations with similar or dis-similar entities. As such, different pretrained embedding such as Word2Vec, GloVe, FastText have been developed. These embeddings generated on millions of words are however very large in terms of size. Having embeddings with floating point precision also makes the downstream evaluation slow. In this paper we present a novel method to convert continuous embedding to its binary representation, thus reducing the overall size of the embedding while keeping the semantic and relational knowledge intact. This will facilitate an option of porting such big embedding onto devices where space is limited. We also present different approaches suitable for different downstream tasks based on the requirement of contextual and semantic information. Experiments have shown comparable result in downstream tasks with 7 to 15 times reduction in file size and about 5% change in evaluation parameters.
Cite
CITATION STYLE
Navali, S., Sherki, P. P., Inturi, R., & Vala, V. (2020). Word Embedding Binarization with Semantic Information Preservation. In COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference (pp. 1256–1265). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.coling-main.108
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.