Abstract
Human action recognition (HAR) is essential in applications ranging from surveillance to interactive gaming. Traditional methods, which typically handle spatial or temporal data separately, often fail to integrate these dimensions effectively. This study introduces the novel DHAR-Net (Deep Hybrid Action Recognition Network) model, a hybrid deep learning framework that integrates convolutional and transformer networks to process spatial and temporal data concurrently, significantly enhancing action recognition accuracy. DHAR-Net was evaluated using the MSR Action3D and UTD-MHAD datasets, which cover a broad range of human actions captured through depth cameras and RGB videos. When compared to conventional models such as LSTM, Edge CNN, and 3D CNN, DHAR-Net demonstrates superior performance, achieving an accuracy of 98.54%, precision of 98.36%, and F1-score of 98.52% on the MSR Action3D dataset, and an accuracy of 98.27%, precision of 97.93%, and F1-score of 98.11% on the UTD-MHAD dataset. This performance demonstrates DHAR-Net’s considerable effectiveness in Human Activity Recognition, indicating substantial advancement in the field.
Author supplied keywords
Cite
CITATION STYLE
Rao, D. S., Ramyasree, K., Ahmad, S. M. K. M. A., & Rao, M. S. (2025). DHAR-Net: A Deep Learning Framework for Hybrid Applications that Combines Motion and Depth Data to Improve Human Action Recognition. International Journal of Intelligent Engineering and Systems, 18(3), 133–149. https://doi.org/10.22266/ijies2025.0430.10
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.