Abstract
The capacity to comprehend and interact with others through language is the most valuable human ability. Since emotions are crucial to communication, we are well-trained to recognize and interpret the many emotions we encounter. Contrary to popular assumption, the subjective aspect of human mood makes emotion recognition difficult for computers. There are some works based on Emotion recognition using images, text, and audio. We are here working on the audio dataset to find the accurate human emotion for computers to understand. In this work, we have utilized a Long Short-Term Memory (LSTM) model to implement Speech Emotion Recognition (SER) from Audio data on two different datasets: the Toronto Emotional Speech Set (TESS) and the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). The accuracy rates of our LSTM-based model were impressive, with 91.25% for the RAVDESS dataset and 98.05% for the TESS dataset; the combined accuracy for both datasets was 87.66%. These results highlight the effectiveness of the LSTM model in effectively identifying and categorizing emotional states from audio files. The study adds significant knowledge to the field of speech emotion recognition by emphasizing the model’s ability to handle a variety of datasets and its potential.
Author supplied keywords
Cite
CITATION STYLE
Or-Rashid, M. M., Nondi, A. K., Sadnun, A. A., Wadud, M. A. H., Bhuiyan, T. M. A. U. H., & Hossain, M. S. (2025). Speech Emotion Recognition from Audio Data Using LSTM Model. International Journal of Advanced Computer Science and Applications, 16(7), 247–254. https://doi.org/10.14569/IJACSA.2025.0160726
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.