Modeling the paraphrase detection task over a heterogeneous graph network with data augmentation

5Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

Paraphrase detection is a Natural-Language Processing (NLP) task that aims at automatically identifying whether two sentences convey the same meaning (even with different words). For the Portuguese language, most of the works model this task as a machine-learning solution, extracting features and training a classifier. In this paper, following a different line, we explore a graph structure representation and model the paraphrase identification task over a heterogeneous network. We also adopt a back-translation strategy for data augmentation to balance the dataset we use. Our approach, although simple, outperforms the best results reported for the paraphrase detection task in Portuguese, showing that graph structures may capture better the semantic relatedness among sentences.

Cite

CITATION STYLE

APA

Anchiêta, R. T., de Sousa, R. F., & Pardo, T. A. S. (2020). Modeling the paraphrase detection task over a heterogeneous graph network with data augmentation. Information (Switzerland), 11(9), 1–12. https://doi.org/10.3390/info11090422

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free