A Sense-Topic Model for Word Sense Induction with Unsupervised Data Enrichment

  • Wang J
  • Bansal M
  • Gimpel K
  • et al.
N/ACitations
Citations of this article
112Readers
Mendeley users who have this article in their library.

Abstract

Word sense induction (WSI) seeks to automatically discover the senses of a word in a corpus via unsupervised methods. We propose a sense-topic model for WSI, which treats sense and topic as two separate latent variables to be inferred jointly. Topics are informed by the entire document, while senses are informed by the local context surrounding the ambiguous word. We also discuss unsupervised ways of enriching the original corpus in order to improve model performance, including using neural word embeddings and external corpora to expand the context of each data instance. We demonstrate significant improvements over the previous state-of-the-art, achieving the best results reported to date on the SemEval-2013 WSI task.

Cite

CITATION STYLE

APA

Wang, J., Bansal, M., Gimpel, K., Ziebart, B. D., & Yu, C. T. (2015). A Sense-Topic Model for Word Sense Induction with Unsupervised Data Enrichment. Transactions of the Association for Computational Linguistics, 3, 59–71. https://doi.org/10.1162/tacl_a_00122

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free