This paper designs a Monolingual Lexicon Induction task and observes that two factors accompany the degraded accuracy of bilingual lexicon induction for rare words. First, a diminishing margin between similarities in low frequency regime, and secondly, exacerbated hubness at low frequency. Based on the observation, we further propose two methods to address these two factors, respectively. The larger issue is hubness. Addressing that improves induction accuracy significantly, especially for low-frequency words.
CITATION STYLE
Huang, J., Cai, X., & Church, K. (2020). Improving bilingual lexicon induction for low frequency words. In EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference (pp. 1310–1314). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.emnlp-main.100
Mendeley helps you to discover research relevant for your work.