View-invariant 3D Skeleton-based Human Activity Recognition based on Transformer and Spatio-temporal Features

11Citations
Citations of this article
4Readers
Mendeley users who have this article in their library.
Get full text

Abstract

With the emergence of depth sensors, real-time 3D human skeleton estimation have become easier to accomplish. Thus, methods for human activity recognition (HAR) based on 3D skeleton have become increasingly accessible. In this paper, we introduce a new approach for human activity recognition using 3D skeletal data. Our approach generates a set of spatio-temporal and view-invariant features from the skeleton joints. Then, the extracted features are analyzed using a typical Transformer encoder in order to recognize the activity. In fact, Transformers, which are based on self-attention mechanism, have been successful in many domains in the last few years, which makes them suitable for HAR. The proposed approach shows promising performance on different well-known datasets that provide 3D skeleton data, namely, KARD, Florence 3D, UTKinect Action 3D and MSR Action 3D.

Cite

CITATION STYLE

APA

Snoun, A., Bouchrika, T., & Jemai, O. (2022). View-invariant 3D Skeleton-based Human Activity Recognition based on Transformer and Spatio-temporal Features. In International Conference on Pattern Recognition Applications and Methods (Vol. 1, pp. 706–715). Science and Technology Publications, Lda. https://doi.org/10.5220/0010895300003122

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free