Hierarchical Multi-modal Transformer for Automatic Detection of COVID-19

8Citations
Citations of this article
13Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Automated COVID-19 detection based on analysis of cough recordings has been an important field of study, as efficient and accurate methods are necessary to contain the spread of the global pandemic and relieve the burden on medical facilities. While previous works presented lightweight machine learning models [9], these models may sacrifice accuracy and interpretability to integrate into mobile devices. Besides, the question of how to effectively associate indicators from audio signals to other modality inputs (i.e. patient information) is still largely unexplored, as previous works predominantly relied on simply concatenated features to learn. To tackle these issues, this paper proposes a novel Hierarchical Multi-modal Transformer (HMT) that learns more informative multi-modal representations with a cross attention module during the feature fusion procedure. Besides, the block aggregation algorithm for the HMT provides an efficient and improved solution from the Vanilla Vision Transformer for limited COVID-19 benchmark datasets. Extensive experiments show the effectiveness of our proposed model for more accurate COVID-19 detection, which yield state-of-the-art results on two public datasets, Coswara and COUGHVID.

Cite

CITATION STYLE

APA

Tang, S., Hu, X., Atlas, L., Khanzada, A., & Pilanci, M. (2022). Hierarchical Multi-modal Transformer for Automatic Detection of COVID-19. In ACM International Conference Proceeding Series (pp. 197–202). Association for Computing Machinery. https://doi.org/10.1145/3556384.3556414

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free