Abstract
Emotion recognition in speech has gained increasing relevance in recent years, enabling more personalized interactions between users and automated systems. This paper presents the development of a dataset of features obtained from RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song) to classify emotions in speech. The paper highlights audio processing techniques such as silence removal and framing to extract features from the recordings. The features are extracted from the audio signals using spectral techniques, time-domain analysis, and the discrete wavelet transform. The resulting dataset is used to train a neural network and the support vector machine learning algorithm. Cross-validation is employed for model training. The developed models were optimized using a software package that performs hyperparameter tuning to improve results. Finally, the emotional classification outcomes were compared. The results showed an emotion classification accuracy of 0.654 for the perceptron neural network and 0.724 for the support vector machine algorithm, demonstrating satisfactory performance in emotion classification.
Author supplied keywords
Cite
CITATION STYLE
Colunga-Rodriguez, A. A., Martínez-Rebollar, A., Estrada-Esquivel, H., Clemente, E., & Pliego-Martínez, O. A. (2025). Developing a Dataset of Audio Features to Classify Emotions in Speech. Computation, 13(2). https://doi.org/10.3390/computation13020039
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.