Improving translation memory matching and retrieval using paraphrases

Rohit Gupta; Constantin Orăsan; Marcos Zampieri; Mihaela Vela; Josef van Genabith; Ruslan Mitkov

Journal Article

Improving translation memory matching and retrieval using paraphrases

Machine Translation (2016) 30(1-2) 19-40

DOI: 10.1007/s10590-016-9180-0

7Citations

15Readers

Get full text

Abstract

Most current translation memory (TM) systems work on the string level (character or word level) and lack semantic knowledge while matching. They use simple edit-distance (ED) calculated on the surface form or some variation on it (stem, lemma), which does not take into consideration any semantic aspects in matching. This paper presents a novel and efficient approach to incorporating semantic information in the form of paraphrasing (PP) in the ED metric. The approach computes ED while efficiently considering paraphrases using dynamic programming and greedy approximation. In addition to using automatic evaluation metrics like BLEU and METEOR, we have carried out an extensive human evaluation in which we measured post-editing time, keystrokes, HTER, HMETEOR, and carried out three rounds of subjective evaluations. Our results show that PP substantially improves TM matching and retrieval, resulting in translation performance increases when translators use paraphrase-enhanced TMs.

Author supplied keywords

Cite

CITATION STYLE

APA

Gupta, R., Orăsan, C., Zampieri, M., Vela, M., van Genabith, J., & Mitkov, R. (2016). Improving translation memory matching and retrieval using paraphrases. Machine Translation, 30(1–2), 19–40. https://doi.org/10.1007/s10590-016-9180-0

Improving translation memory matching and retrieval using paraphrases

Abstract

Author supplied keywords

Cite

Register to see more suggestions