Improving Vietnamese-English Cross-Lingual Retrieval for Legal and General Domains

7Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Document retrieval plays a crucial role in numerous question-answering systems, yet research has concentrated on the general knowledge domain and resource-rich languages like English. In contrast, it remains largely underexplored in low-resource languages and cross-lingual scenarios within specialized domain knowledge such as legal. We present a novel dataset designed for cross-lingual retrieval between Vietnamese and English, which not only covers the general domain but also extends to the legal field. Additionally, we propose auxiliary loss function and symmetrical training strategy that significantly enhance the performance of state-of-the-art models on these retrieval tasks. Our contributions offer a significant resource and methodology aimed at improving cross-lingual retrieval in both legal and general QA settings, facilitating further advancements in document retrieval research across multiple languages and a broader spectrum of specialized domains. All the resources related to our work can be accessed at huggingface.co/datasets/ bkai-foundation-models/crosslingual.

Cite

CITATION STYLE

APA

Nguyen, T. N., Le Hai, N., Hieu, N. D., Nguyen, D. A., Van, L. N., Nguyen, T. H., & Dinh, S. (2025). Improving Vietnamese-English Cross-Lingual Retrieval for Legal and General Domains. In Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025 (Vol. 2, pp. 142–153). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.naacl-short.12

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free