Abstract
There is a rising interest in dynamic gesture recognition as a research area. This is the result of emerging global pandemics as well as the need to avoid touching different surfaces. Most of the previous research has focused on implementing deep learning algorithms for the RGB modality. However, despite its potential to enhance the algorithm’s performance, gesture recognition has not widely utilised the concept of attention. Most research also used three-dimensional convolutional networks with long short-term memory networks for gesture recognition. However, these networks can be computationally expensive. As a result, this paper employs pre-trained models in conjunction with the skeleton modality to address the challenges posed by background noise. The goal is to present a comparative analysis of various gesture recognition models, divided based on video frames or skeletons. The performance of different models was evaluated using a dataset taken from Kaggle with a size of 2 GB. Each video contains 30 frames (or images) to recognise five gestures. The transformer model for skeleton-based gesture recognition achieves 0.99 accuracy and can be used to capture temporal dependencies in sequential data.
Author supplied keywords
Cite
CITATION STYLE
Althubiti, A. H., & Algethami, H. (2024). Dynamic Gesture Recognition using a Transformer and Mediapipe. International Journal of Advanced Computer Science and Applications, 15(6), 1424–1439. https://doi.org/10.14569/IJACSA.2024.01506143
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.