Dynamic Gesture Recognition using a Transformer and Mediapipe

4Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.
Get full text

Abstract

There is a rising interest in dynamic gesture recognition as a research area. This is the result of emerging global pandemics as well as the need to avoid touching different surfaces. Most of the previous research has focused on implementing deep learning algorithms for the RGB modality. However, despite its potential to enhance the algorithm’s performance, gesture recognition has not widely utilised the concept of attention. Most research also used three-dimensional convolutional networks with long short-term memory networks for gesture recognition. However, these networks can be computationally expensive. As a result, this paper employs pre-trained models in conjunction with the skeleton modality to address the challenges posed by background noise. The goal is to present a comparative analysis of various gesture recognition models, divided based on video frames or skeletons. The performance of different models was evaluated using a dataset taken from Kaggle with a size of 2 GB. Each video contains 30 frames (or images) to recognise five gestures. The transformer model for skeleton-based gesture recognition achieves 0.99 accuracy and can be used to capture temporal dependencies in sequential data.

Cite

CITATION STYLE

APA

Althubiti, A. H., & Algethami, H. (2024). Dynamic Gesture Recognition using a Transformer and Mediapipe. International Journal of Advanced Computer Science and Applications, 15(6), 1424–1439. https://doi.org/10.14569/IJACSA.2024.01506143

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free