MP-GestLSTM: real time gesture detection using MediaPipe and LSTM

4Citations
Citations of this article
20Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Sign language is a crucial form of communication for individuals with hearing impairments; thus, the development of accurate and efficient sign language recognition systems is of significant importance. The main aim of the work is to develop an accurate and efficient sign language recognition model using MediaPipe (MP) and deep learning techniques. MediaPipe (BlazePose) is a top-down model with an encoder-decoder architecture to predict the keypoints for all joints. MediaPipe detects 33 body(pose) landmarks, 21 hand landmarks, and 478 3-dimensional face landmarks. A custom dataset called MP-Gest consists of 800 videos belonging to 20 classes at the word level of American Sign Language (ASL) is generated by two signers. A selection of MediaPipe keypoints is extracted from this custom dataset and used to train a Long Short-Term Memory (LSTM)-based custom model called MP-GestLSTM. The MP-GestLSTM model achieved a validation accuracy of 93.75% and testing accuracy of 94.38%, indicating that the model outperformed existing research with minimal error. The proposed model gives enhanced recognition accuracy using a simple custom architecture and contributes to the advancement of sign language recognition by overcoming the existing limitations.

Cite

CITATION STYLE

APA

Varshini, T. S., & Rukmani, P. (2025). MP-GestLSTM: real time gesture detection using MediaPipe and LSTM. Systems Science and Control Engineering, 13(1). https://doi.org/10.1080/21642583.2025.2587853

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free