Abstract
Isolated Sign Language Recognition (ISLR), which focuses on identifying individual signs from sign language videos, presents substantial challenges due to small and ambiguous hand regions, high visual similarity among signs, and large intra-class variability. This study investigates the adaptability of YOLO-Act, a unified spatiotemporal detection framework originally developed for generic action recognition in videos, when applied to large-scale sign language benchmarks. YOLO-Act jointly performs signer localization (identifying the person signing within a video) and action classification (determining which sign is performed) directly from RGB sequences, eliminating the need for pose estimation or handcrafted temporal cues. We evaluate the model on the WLASL2000 and MSASL1000 datasets for American Sign Language recognition, achieving Top-1 accuracies of 67.07% and 81.41%, respectively. The latter represents a 3.55% absolute improvement over the best-performing baseline without pose supervision. These results demonstrate the strong cross-domain generalization and robustness of YOLO-Act in complex multi-class recognition scenarios.
Author supplied keywords
Cite
CITATION STYLE
Alzahrani, N., Bchir, O., & Ben Ismail, M. M. (2025). Unified Spatiotemporal Detection for Isolated Sign Language Recognition Using YOLO-Act. Electronics (Switzerland), 14(23). https://doi.org/10.3390/electronics14234589
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.