Audio-Visual Action Recognition Using Transformer Fusion Network

10Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Our approach to action recognition is grounded in the intrinsic coexistence of and complementary relationship between audio and visual information in videos. Going beyond the traditional emphasis on visual features, we propose a transformer-based network that integrates both audio and visual data as inputs. This network is designed to accept and process spatial, temporal, and audio modalities. Features from each modality are extracted using a single Swin Transformer, originally devised for still images. Subsequently, these extracted features from spatial, temporal, and audio data are adeptly combined using a novel modal fusion module (MFM). Our transformer-based network effectively fuses these three modalities, resulting in a robust solution for action recognition.

Cite

CITATION STYLE

APA

Kim, J. H., & Won, C. S. (2024). Audio-Visual Action Recognition Using Transformer Fusion Network. Applied Sciences (Switzerland), 14(3). https://doi.org/10.3390/APP14031190

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free