Speech Emotion Recognition using Threshold Fusion for Enhancing Audio Sensitivity

0Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Speech Emotion Recognition (SER) has found applications in various fields. However, most SER studies exhibit a bias towards the text modality, which can lead to incorrect recognition when nonverbal audio features convey the primary emotional information. To address this issue, we propose a two-step solution to enhance the audio emotion sensitivity of SER models. First, we use a parallel emotional speech dataset (ESD), which contains identical speech content pronounced with different emotions, to pretrain a speech content-independent emotion recognition model, named the Audio Sensitive Network (ASN). Second, we propose a novel threshold fusion technique utilizing the Tree-structured Parzen Estimator (TPE) to optimize different thresholds for each predictive label, integrating the ASN with baseline SER classifiers. To demonstrate the efficacy of our approach, we conduct experiments on the IEMOCAP and ESD datasets. The results reveal that our novel method enhances audio sensitivity by enhancing the performance of existing SER classifiers.

Cite

CITATION STYLE

APA

Luo, Z., Christiansson, S., Ladóczki, B., & Komatani, K. (2023). Speech Emotion Recognition using Threshold Fusion for Enhancing Audio Sensitivity. In Proceedings of the 5th ACM International Conference on Multimedia in Asia, MMAsia 2023 Workshops. Association for Computing Machinery, Inc. https://doi.org/10.1145/3611380.3628557

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free