Developing corpora using word2vec and wikipedia for word sense disambiguation

11Citations
Citations of this article
25Readers
Mendeley users who have this article in their library.

Abstract

Word Sense Disambiguation (WSD) is one of the most difficult problems in the artificial intelligence field or well known as AI-hard or AI-complete. A lot of problems can be solved using word sense disambiguation approach such as sentiment analysis, machine translation, search engine relevance, coherence, anaphora resolution, and inference. This research is done to solve WSD problem with two small corpora. The use of Word2vec and Wikipedia are proposed to develop the corpora. After developing the corpora, the similarity of the sentence with the corpora is measured using cosine similarity to determine the meaning of the ambiguous word. Lastly, to improve accuracy, Lesk algorithms and Wu Palmer similarity are used to deal with problems when there is no word from a sentence in the corpus. The results of the research show an 85.51% accuracy rate and the semantic similarity improve the accuracy rate by 8.02% in determining the meaning of ambiguous words.

Cite

CITATION STYLE

APA

Nurifan, F., Sarno, R., & Wahyuni, C. S. (2018). Developing corpora using word2vec and wikipedia for word sense disambiguation. Indonesian Journal of Electrical Engineering and Computer Science, 12(3), 1239–1246. https://doi.org/10.11591/ijeecs.v12.i3.pp1239-1246

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free