Cross-modal Contrastive Learning for Speech Translation

85Citations
Citations of this article
62Readers
Mendeley users who have this article in their library.

Abstract

How can we learn unified representations for spoken utterances and their written text? Learning similar representations for semantically similar speech and text is important for speech translation. To this end, we propose ConST, a cross-modal contrastive learning method for end-to-end speech-to-text translation. We evaluate ConST and a variety of previous baselines on a popular benchmark MuST-C. Experiments show that the proposed ConST consistently outperforms the previous methods, and achieves an average BLEU of 29.4. The analysis further verifies that ConST indeed closes the representation gap of different modalities - its learned representation improves the accuracy of cross-modal speech-text retrieval from 4% to 88%. Code and models are available at https://github.com/ReneeYe/ConST.

Cite

CITATION STYLE

APA

Ye, R., Wang, M., & Li, L. (2022). Cross-modal Contrastive Learning for Speech Translation. In NAACL 2022 - 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference (pp. 5099–5113). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.naacl-main.376

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free