Word-like character n-gram embedding

5Citations
Citations of this article
82Readers
Mendeley users who have this article in their library.

Abstract

We propose a new word embedding method called word-like character n-gram embedding, which learns distributed representations of words by embedding word-like character ngrams. Our method is an extension of recently proposed segmentation-free word embedding, which directly embeds frequent character ngrams from a raw corpus. However, its n-gram vocabulary tends to contain too many non-word n-grams. We solved this problem by introducing an idea of expected word frequency. Compared to the previously proposed methods, our method can embed more words, along with the words that are not included in a given basic word dictionary. Since our method does not rely on word segmentation with rich word dictionaries, it is especially effective when the text in the corpus is in unsegmented language and contains many neologisms and informal words (e.g., Chinese SNS dataset). Our experimental results on Sina Weibo (a Chinese microblog service) and Twitter show that the proposed method can embed more words and improve the performance of downstream tasks.

Cite

CITATION STYLE

APA

Kim, G., Fukui, K., & Shimodaira, H. (2018). Word-like character n-gram embedding. In 4th Workshop on Noisy User-Generated Text, W-NUT 2018 - Proceedings of the Workshop (pp. 148–152). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-6120

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free