Multimodal Speech Emotion Recognition Based on Large Language Model

N/ACitations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

Currently, an increasing number of tasks in speech emotion recognition rely on the analysis of both speech and text features. However, there remains a paucity of research exploring the potential of leveraging large language models like GPT-3 to enhance emotion recognition. In this investigation, we harness the power of the GPT-3 model to extract semantic information from transcribed texts, generating text modal features with a dimensionality of 1536. Subsequently, we perform feature fusion, combining the 1536-dimensional text features with 1188-dimensional acoustic features to yield comprehensive multi-modal recognition outcomes. Our findings reveal that the proposed method achieves a weighted accuracy of 79.62% across the four emotion categories in IEMOCAP, underscoring the considerable enhancement in emotion recognition accuracy facilitated by integrating large language models.

Cite

CITATION STYLE

APA

Congcong, F., Yun, J., Guanlin, C., Yunfan, Z., Shidang, L., Ma, Y., & Yue, X. (2024). Multimodal Speech Emotion Recognition Based on Large Language Model. IEICE Transactions on Information and Systems, E107.D(11), 1463–1467. https://doi.org/10.1587/transinf.2024EDL8034

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free