Speech Emotion Recognition from Audio Data Using LSTM Model

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

The capacity to comprehend and interact with others through language is the most valuable human ability. Since emotions are crucial to communication, we are well-trained to recognize and interpret the many emotions we encounter. Contrary to popular assumption, the subjective aspect of human mood makes emotion recognition difficult for computers. There are some works based on Emotion recognition using images, text, and audio. We are here working on the audio dataset to find the accurate human emotion for computers to understand. In this work, we have utilized a Long Short-Term Memory (LSTM) model to implement Speech Emotion Recognition (SER) from Audio data on two different datasets: the Toronto Emotional Speech Set (TESS) and the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). The accuracy rates of our LSTM-based model were impressive, with 91.25% for the RAVDESS dataset and 98.05% for the TESS dataset; the combined accuracy for both datasets was 87.66%. These results highlight the effectiveness of the LSTM model in effectively identifying and categorizing emotional states from audio files. The study adds significant knowledge to the field of speech emotion recognition by emphasizing the model’s ability to handle a variety of datasets and its potential.

Cite

CITATION STYLE

APA

Or-Rashid, M. M., Nondi, A. K., Sadnun, A. A., Wadud, M. A. H., Bhuiyan, T. M. A. U. H., & Hossain, M. S. (2025). Speech Emotion Recognition from Audio Data Using LSTM Model. International Journal of Advanced Computer Science and Applications, 16(7), 247–254. https://doi.org/10.14569/IJACSA.2025.0160726

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free