Challenging the Boundaries of Unsupervised Learning for Semantic Similarity

38Citations
Citations of this article
49Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

The semantic analysis field has a crucial role to play in the research related to text analytics. Calculating the semantic similarity between sentences is a long-standing problem in the area of natural language processing, and it differs significantly as the domain of operation differs. In this paper, we present a methodology that can be applied across multiple domains by incorporating corpora-based statistics into a standardized semantic similarity algorithm. To calculate the semantic similarity between words and sentences, the proposed method follows an edge-based approach using a lexical database. When tested on both benchmark standards and mean human similarity dataset, the methodology achieves a high correlation value for both word ( r=0.8753 ) and sentence similarity ( r=0.8793 ) concerning Rubenstein and Goodenough standard and the SICK dataset ( r=0.8324 1 ) outperforming other unsupervised models. 1 Eliminating the outliers which constitutes to 3.75% of 4927 statement pairs.

Cite

CITATION STYLE

APA

Pawar, A., & Mago, V. (2019). Challenging the Boundaries of Unsupervised Learning for Semantic Similarity. IEEE Access, 7, 16291–16308. https://doi.org/10.1109/ACCESS.2019.2891692

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free