Abstract
Automated COVID-19 detection based on analysis of cough recordings has been an important field of study, as efficient and accurate methods are necessary to contain the spread of the global pandemic and relieve the burden on medical facilities. While previous works presented lightweight machine learning models [9], these models may sacrifice accuracy and interpretability to integrate into mobile devices. Besides, the question of how to effectively associate indicators from audio signals to other modality inputs (i.e. patient information) is still largely unexplored, as previous works predominantly relied on simply concatenated features to learn. To tackle these issues, this paper proposes a novel Hierarchical Multi-modal Transformer (HMT) that learns more informative multi-modal representations with a cross attention module during the feature fusion procedure. Besides, the block aggregation algorithm for the HMT provides an efficient and improved solution from the Vanilla Vision Transformer for limited COVID-19 benchmark datasets. Extensive experiments show the effectiveness of our proposed model for more accurate COVID-19 detection, which yield state-of-the-art results on two public datasets, Coswara and COUGHVID.
Author supplied keywords
Cite
CITATION STYLE
Tang, S., Hu, X., Atlas, L., Khanzada, A., & Pilanci, M. (2022). Hierarchical Multi-modal Transformer for Automatic Detection of COVID-19. In ACM International Conference Proceeding Series (pp. 197–202). Association for Computing Machinery. https://doi.org/10.1145/3556384.3556414
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.