Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Sign Language Translation (SLT) aims to convert sign language videos into spoken or written text. While early systems relied on gloss annotations as an intermediate supervision, such annotations are costly to obtain and often fail to capture the full complexity of continuous signing. In this work, we propose a two-phase, dual visual encoder framework for gloss-free SLT, leveraging contrastive visual-language pretraining. During pretraining, our approach employs two complementary visual backbones whose outputs are jointly aligned with each other and with sentence-level text embeddings via a contrastive objective. During the downstream SLT task, we fuse the visual features and input them into an encoder-decoder model. On the Phoenix-2014T benchmark, our dual encoder architecture consistently outperforms its single-stream variants and achieves the highest BLEU-4 score among existing gloss-free SLT approaches.

Cite

CITATION STYLE

APA

Mercanoglu Sincan, O., & Bowden, R. (2025). Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation. In IVA 2025 - Adjunct Proceedings of the 25th ACM International Conference on Intelligent Virtual Agents. Association for Computing Machinery, Inc. https://doi.org/10.1145/3742886.3756703

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free