Abstract
In this paper, we describe the TATO system which participated in the SemEval-2015 Task 2a: “Semantic Textual Similarity (STS) for English”. Our system is trained on published datasets from the previous competitions. Based on some machine learning techniques, it combines multiple similarity measures of varying complexity ranging from simple lexical and syntactic similarity measures to complex semantic similarity ones to compute semantic textual similarity. Our final model consists of a simple linear combination of about 30 main features out of a numerous number of features experimented. The results are promising, with Pearson's coefficients on each individual dataset ranging from 0.6796 to 0.8167 and an overall weighted mean score of 0.7422, well above the task baseline system.
Cite
CITATION STYLE
Vu, T. T., Tran, Q. H., & Pham, S. B. (2015). TATO: Leveraging on Multiple Strategies for Semantic Textual Similarity. In SemEval 2015 - 9th International Workshop on Semantic Evaluation, co-located with the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2015 - Proceedings (pp. 190–195). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/s15-2034
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.