MVDT: Multiview Distillation Transformer for View-Invariant Sign Language Translation

N/ACitations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

Sign language translation based on machine learning plays a crucial role in facilitating communication between deaf and hearing individuals. However, due to the complexity and variability of sign language, coupled with limited observation angles, single-view sign language translation models often underperform in real-world applications. Although some studies have attempted to improve translation efficiency by incorporating multiview data, challenges, such as feature alignment, fusion, and the high cost of capturing multiview data, remain significant barriers in many practical scenarios. To address these issues, we propose a multiview distillation transformer model (MVDT) for continuous sign language translation. The MVDT introduces a novel distillation mechanism, where a teacher model is designed to learn common features from multiview data, subsequently guiding a student model to extract view-invariant features using only single-view input. To evaluate the proposed method, we construct a multiview sign language dataset comprising five distinct views and conduct extensive experiments comparing the MVDT with state-of-the-art methods. Experimental results demonstrate that the proposed model exhibits superior view-invariant translation capabilities across different views.

Cite

CITATION STYLE

APA

Guan, Z., Hu, Y., Jiang, H., Sun, Y., & Yin, B. (2025). MVDT: Multiview Distillation Transformer for View-Invariant Sign Language Translation. IET Computer Vision, 19(1). https://doi.org/10.1049/cvi2.70038

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free