Abstract
This study explores methods for developing Machine Translation dictionaries on the basis of word frequency lists coming from comparable corpora. We investigate (1) various methods to measure the similarity of cognates between related languages, (2) detection and removal of noisy cognate translations using SVM ranking. We show preliminary results on several Romance and Slavonic languages.
Cite
CITATION STYLE
Rios, M., & Sharoff, S. (2015). Obtaining SMT dictionaries for related languages. In 8th Workshop on Building and Using Comparable Corpora, BUCC 2015 - co-located with 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing, ACL-IJCNLP 2015 - Proceedings (pp. 68–73). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w15-3410
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.