A Generative Model for Extracting Parallel Fragments from Comparable Documents

4Citations
Citations of this article
76Readers
Mendeley users who have this article in their library.

Abstract

Although parallel corpora are essential language resources for many NLP tasks, they are rare or even not available for many language pairs. Instead, comparable corpora are widely available and contain parallel fragments of information that can be used applications like statistical machine translations. In this research, we propose a generative LDA based model for extracting parallel fragments from comparable documents without using any initial parallel data or bilingual lexicon. The experimental results show significant improvement if the extracted sentence fragments generated by the proposed method are used in addition to an existing parallel corpus in an SMT task. According to human judgment, the accuracy of the proposed method for an English-Persian task is about 66%. Also, the OOV rate for the same task is reduced by 28%.

Cite

CITATION STYLE

APA

Bakhshaei, S., Khadivi, S., & Safabakhsh, R. (2015). A Generative Model for Extracting Parallel Fragments from Comparable Documents. In 8th Workshop on Building and Using Comparable Corpora, BUCC 2015 - co-located with 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing, ACL-IJCNLP 2015 - Proceedings (pp. 43–51). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w15-3407

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free