A Simple and Effective Usage of Word Clusters for CBOW Model

  • Feng Y
  • Hu C
  • Kamigaito H
  • et al.
N/ACitations
Citations of this article
65Readers
Mendeley users who have this article in their library.

Abstract

We propose a simple and effective method for incorporating word clusters into the Continuous Bag-of-Words (CBOW) model. Specifically, we propose to replace infrequent input and output words in CBOW model with their clusters. The resulting cluster-incorporated CBOW model produces embeddings of frequent words and a small amount of cluster embeddings, which will be fine-tuned in downstream tasks. We empirically show our replacing method works well on several downstream tasks. Through our analysis, we show that our method might be also useful for other similar models which produce word embeddings.

Cite

CITATION STYLE

APA

Feng, Y., Hu, C., Kamigaito, H., Takamura, H., & Okumura, M. (2022). A Simple and Effective Usage of Word Clusters for CBOW Model. Journal of Natural Language Processing, 29(3), 785–806. https://doi.org/10.5715/jnlp.29.785

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free