Developing a Dataset of Audio Features to Classify Emotions in Speech

10Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

Emotion recognition in speech has gained increasing relevance in recent years, enabling more personalized interactions between users and automated systems. This paper presents the development of a dataset of features obtained from RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song) to classify emotions in speech. The paper highlights audio processing techniques such as silence removal and framing to extract features from the recordings. The features are extracted from the audio signals using spectral techniques, time-domain analysis, and the discrete wavelet transform. The resulting dataset is used to train a neural network and the support vector machine learning algorithm. Cross-validation is employed for model training. The developed models were optimized using a software package that performs hyperparameter tuning to improve results. Finally, the emotional classification outcomes were compared. The results showed an emotion classification accuracy of 0.654 for the perceptron neural network and 0.724 for the support vector machine algorithm, demonstrating satisfactory performance in emotion classification.

Cite

CITATION STYLE

APA

Colunga-Rodriguez, A. A., Martínez-Rebollar, A., Estrada-Esquivel, H., Clemente, E., & Pliego-Martínez, O. A. (2025). Developing a Dataset of Audio Features to Classify Emotions in Speech. Computation, 13(2). https://doi.org/10.3390/computation13020039

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free