Compiling a massive, multilingual dictionary via probabilistic inference

undefined Mausam; Stephen Soderland; Oren Etzioni; Daniel S. Weld; Michael Skinner; Jeff Bilmes

Conference ProceedingsOPEN ACCESS

Compiling a massive, multilingual dictionary via probabilistic inference

ACL-IJCNLP 2009 - Joint Conf. of the 47th Annual Meeting of the Association for Computational Linguistics and 4th Int. Joint Conf. on Natural Language Processing of the AFNLP, Proceedings of the Conf. (2009) 262-270

DOI: 10.3115/1687878.1687917

54Citations

118Readers

Abstract

Can we automatically compose a large set of Wiktionaries and translation dictionaries to yield a massive, multilingual dictionary whose coverage is substantially greater than that of any of its constituent dictionaries? The composition of multiple translation dictionaries leads to a transitive inference problem: if word A translates to word B which in turn translates to word C, what is the probability that C is a translation of A? The paper introduces a novel algorithm that solves this problem for 10,000,000 words in more than 1,000 languages. The algorithm yields PANDICTIONARY, a novel multilingual dictionary. PANDICTIONARY contains more than four times as many translations than in the largest Wiktionary at precision 0.90 and over 200,000,000 pairwise translations in over 200,000 language pairs at precision 0.8. © 2009 ACL and AFNLP.

Cite

CITATION STYLE

APA

Mausam, Soderland, S., Etzioni, O., Weld, D. S., Skinner, M., & Bilmes, J. (2009). Compiling a massive, multilingual dictionary via probabilistic inference. In ACL-IJCNLP 2009 - Joint Conf. of the 47th Annual Meeting of the Association for Computational Linguistics and 4th Int. Joint Conf. on Natural Language Processing of the AFNLP, Proceedings of the Conf. (pp. 262–270). Association for Computational Linguistics (ACL). https://doi.org/10.3115/1687878.1687917

Compiling a massive, multilingual dictionary via probabilistic inference

Abstract

Cite

Register to see more suggestions