Abstract
Speech emotion recognition (SER) remains a challenging task due to the limited affective cues in unimodal representations and the difficulty of aligning heterogeneous features in multimodal systems. Although multimodal large language models (MLLMs) have recently achieved notable progress in affective understanding, they still suffer from hallucination, weak emotion discrimination, and instability during optimization. To address these challenges, this paper presents EmoBridge, a novel multimodal affective learning framework that integrates a large language model with a speech-aware cross-modal encoder, termed EmoBridge-Former, which extends the Q-Former architecture to model fine-grained audio–text emotional dependencies. In contrast to existing methods, EmoBridge introduces two improved optimization objectives: 1) a Prototypical InfoNCE loss, which enhances cross-modal alignment by pulling samples toward their emotion prototypes and enlarging inter-class separability; and 2) a Dynamic Focal Loss, which adaptively adjusts the weighting of hard and easy emotion samples to mitigate class imbalance and gradient domination. In addition, a soft-prompt fusion mechanism is employed to inject multimodal embeddings into the language model in a parameter-efficient manner, thereby enabling end-to-end affective reasoning. Extensive experiments on the IEMOCAP and MELD benchmarks demonstrate that EmoBridge consistently outperforms state-of-the-art SER models in both accuracy and robustness. The proposed framework provides a unified and scalable paradigm for emotion-aware multimodal representation learning and has the potential to advance real-world speech understanding applications.
Author supplied keywords
Cite
CITATION STYLE
Sun, Y., Zhang, X., & Sun, Y. (2025). EmoBridge: Aligning Speech and Language for Emotion Recognition via Q-Former. IEEE Access, 13, 205601–205611. https://doi.org/10.1109/ACCESS.2025.3636123
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.