Abstract
Sign language recognition faces critical challenges including temporal inconsistencies, inadequate cross-cultural feature representation, and limited real-time adaptability with poor generalization across diverse signing styles. We propose Spatiotemporal Transformer Reinforce Epsilon Greedy (STTREG), a novel architecture that integrates epsilon-guided exploration strategies within spatiotemporal transformer frameworks for efficient multilingual sign language recognition. The key innovation extends discrete epsilon-greedy algorithms to continuous spatiotemporal modeling while maintaining gradient compatibility and optimizing computational efficiency. STTREG manages dual memory states through attention-modulated gates, enabling robust cross-lingual generalization across Indian, American, and Chinese sign languages without separate retraining. The framework demonstrated superior generalization capabilities across diverse signers and environmental conditions while achieving low-latency inference suitable for real-time applications. Comprehensive experiments on three benchmark datasets with over 50,000 samples demonstrate average recognition accuracy of 97.1%, outperforming state-of-the-art methods by 4.2% with 65% reduced inference latency. STTREG establishes theoretical foundations for continuous reinforcement learning in spatiotemporal modeling, while delivering practical advances in generalizable, low-latency multilingual sign language interpretation.
Author supplied keywords
Cite
CITATION STYLE
Kar, H., & Viswanathan, P. (2025). Epsilon-Guided Spatiotemporal Transformer: An Exponential Error Reduction for Multi-Memory Multilingual Sign Interpreter. IEEE Open Journal of the Computer Society, 6, 1884–1895. https://doi.org/10.1109/OJCS.2025.3629116
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.