Abstract
Speech Emotion Recognition (SER) has found applications in various fields. However, most SER studies exhibit a bias towards the text modality, which can lead to incorrect recognition when nonverbal audio features convey the primary emotional information. To address this issue, we propose a two-step solution to enhance the audio emotion sensitivity of SER models. First, we use a parallel emotional speech dataset (ESD), which contains identical speech content pronounced with different emotions, to pretrain a speech content-independent emotion recognition model, named the Audio Sensitive Network (ASN). Second, we propose a novel threshold fusion technique utilizing the Tree-structured Parzen Estimator (TPE) to optimize different thresholds for each predictive label, integrating the ASN with baseline SER classifiers. To demonstrate the efficacy of our approach, we conduct experiments on the IEMOCAP and ESD datasets. The results reveal that our novel method enhances audio sensitivity by enhancing the performance of existing SER classifiers.
Author supplied keywords
Cite
CITATION STYLE
Luo, Z., Christiansson, S., Ladóczki, B., & Komatani, K. (2023). Speech Emotion Recognition using Threshold Fusion for Enhancing Audio Sensitivity. In Proceedings of the 5th ACM International Conference on Multimedia in Asia, MMAsia 2023 Workshops. Association for Computing Machinery, Inc. https://doi.org/10.1145/3611380.3628557
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.