Aiming beyond the obvious: Identifying non-obvious cases in semantic similarity datasets

7Citations
Citations of this article
116Readers
Mendeley users who have this article in their library.

Abstract

Existing datasets for scoring text pairs in terms of semantic similarity contain instances whose resolution differs according to the degree of difficulty. This paper proposes to distinguish obvious from non-obvious text pairs based on superficial lexical overlap and ground-truth labels. We characterise existing datasets in terms of containing difficult cases and find that recently proposed models struggle to capture the non-obvious cases of semantic similarity. We describe metrics that emphasise cases of similarity which require more complex inference and propose that these are used for evaluating systems for semantic similarity.

Cite

CITATION STYLE

APA

Peinelt, N., Liakata, M., & Nguyen, D. (2020). Aiming beyond the obvious: Identifying non-obvious cases in semantic similarity datasets. In ACL 2019 - 57th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (pp. 2792–2798). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p19-1268

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free