Abstract
Document retrieval plays a crucial role in numerous question-answering systems, yet research has concentrated on the general knowledge domain and resource-rich languages like English. In contrast, it remains largely underexplored in low-resource languages and cross-lingual scenarios within specialized domain knowledge such as legal. We present a novel dataset designed for cross-lingual retrieval between Vietnamese and English, which not only covers the general domain but also extends to the legal field. Additionally, we propose auxiliary loss function and symmetrical training strategy that significantly enhance the performance of state-of-the-art models on these retrieval tasks. Our contributions offer a significant resource and methodology aimed at improving cross-lingual retrieval in both legal and general QA settings, facilitating further advancements in document retrieval research across multiple languages and a broader spectrum of specialized domains. All the resources related to our work can be accessed at huggingface.co/datasets/ bkai-foundation-models/crosslingual.
Cite
CITATION STYLE
Nguyen, T. N., Le Hai, N., Hieu, N. D., Nguyen, D. A., Van, L. N., Nguyen, T. H., & Dinh, S. (2025). Improving Vietnamese-English Cross-Lingual Retrieval for Legal and General Domains. In Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025 (Vol. 2, pp. 142–153). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.naacl-short.12
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.