Hubert-LSTM: A Hybrid Model for Artificial Intelligence and Human Speech

  • Baias A
N/ACitations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Speech emotion recognition (SER) is a critical component of human-computer interaction, facilitating seamless communication between individuals and machines. In this paper, we propose a hybrid model, integrating Hubert, a cutting-edge speech recognition model, with LSTM (Long Short-Term Memory), known for its effectiveness in sequence modeling tasks, to enhance emotion recognition accuracy in speech audio files. We explore the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) for our investigation, drawn by its complexity and open accessibility. Our hybrid model combines the semantic features extracted by Hubert with LSTM’s ability to capture temporal relationships in audio sequences, thereby improving emotion recognition performance. Through rigorous experimentation and evaluation on a subset of actors from the RAVDESS dataset, our model achieved promising results, outperforming existing approaches, with a maximum accuracy of 89.1 %.

Cite

CITATION STYLE

APA

Baias, A.-C. (2024). Hubert-LSTM: A Hybrid Model for Artificial Intelligence and Human Speech. Engineering World, 6, 159–169. https://doi.org/10.37394/232025.2024.6.17

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free