Deep Multimodal Fusion of Visual and Auditory Features for Robust Material Recognition

6Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

This paper presents a deep neural network incorporating visual and auditory data fusion to enhance material recognition performance. Traditional recognition techniques relying on single data modalities face accuracy and robustness limitations, especially in complex real-world environments. To address these challenges, we develop a multimodal fusion-based model. The proposed approach first extracts features from input images and sounds separately using CNNs and spectral analysis. A concatenation layer then integrates the visual and auditory features. Extensive experiments demonstrate superior material classification over uni-modal methods, with 100% test accuracy across seven material types. The multi-modal fusion model also demonstrates stronger resilience to noise and illumination variations. This research provides a valuable foundation for robust material perception in intelligent systems.

Cite

CITATION STYLE

APA

Shi, Y., Ong, H. R., Yang, S., & Fan, Y. (2024). Deep Multimodal Fusion of Visual and Auditory Features for Robust Material Recognition. International Journal of Computers, Communications and Control, 19(5), 1–17. https://doi.org/10.15837/ijccc.2024.5.6457

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free