Abstract
Sign language translation based on machine learning plays a crucial role in facilitating communication between deaf and hearing individuals. However, due to the complexity and variability of sign language, coupled with limited observation angles, single-view sign language translation models often underperform in real-world applications. Although some studies have attempted to improve translation efficiency by incorporating multiview data, challenges, such as feature alignment, fusion, and the high cost of capturing multiview data, remain significant barriers in many practical scenarios. To address these issues, we propose a multiview distillation transformer model (MVDT) for continuous sign language translation. The MVDT introduces a novel distillation mechanism, where a teacher model is designed to learn common features from multiview data, subsequently guiding a student model to extract view-invariant features using only single-view input. To evaluate the proposed method, we construct a multiview sign language dataset comprising five distinct views and conduct extensive experiments comparing the MVDT with state-of-the-art methods. Experimental results demonstrate that the proposed model exhibits superior view-invariant translation capabilities across different views.
Author supplied keywords
Cite
CITATION STYLE
Guan, Z., Hu, Y., Jiang, H., Sun, Y., & Yin, B. (2025). MVDT: Multiview Distillation Transformer for View-Invariant Sign Language Translation. IET Computer Vision, 19(1). https://doi.org/10.1049/cvi2.70038
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.