A Study of Methods for the Generation of Domain-Aware Word Embeddings

3Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Word embeddings are essential components for many text data applications. In most work, "out-of-the-box" embeddings trained on general text corpora are used, but they can be less effective when applied to domain-specific settings. Thus, how to create "domain-aware" word embeddings is an interesting open research question. In this paper, we study three methods for creating domain-aware word embeddings based on both general and domain-specific text corpora, including concatenation of embedding vectors, weighted fusion of text data, and interpolation of aligned embedding vectors. Even though the investigated strategies are tailored for domain-specific tasks, they are general enough to be applied to any domain and are not specific to a single task. Experimental results show that all three methods can work well, however, the interpolation method consistently works best.

Cite

CITATION STYLE

APA

Seyler, D., & Zhai, C. X. (2020). A Study of Methods for the Generation of Domain-Aware Word Embeddings. In SIGIR 2020 - Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 1609–1612). Association for Computing Machinery, Inc. https://doi.org/10.1145/3397271.3401287

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free