Abstract
Speech emotion recognition is a challenging task and feature extraction plays an important role in effectively classifying speech into different emotions. In this paper, we apply traditional feature extraction methods like MFCC for feature extraction from audio files. Instead of using traditional machine learning approaches like SVM to classify audio files, we investigate different neural network architectures. Our baseline model implemented as a convolutional neural network results in 60% classification accuracy. We propose a hybrid neural network architecture based on Convolutional and Long Short-Term Memory (ConvLSTM) networks to capture spatial and sequential information of audio files. Our experimental results show that our ComvLSTM model has achieved an accuracy of 59%. We improved our model with data augmentation techniques and re-trained it with augmented dataset. The classification accuracy achieves 91% for multi-class classification of RAVDESS dataset outperforming the accuracy of state-of-the-art multi-class classification models that used the similar data.
Author supplied keywords
Cite
CITATION STYLE
Badr, Y., Mukherjee, P., & Thumati, S. M. (2021). Speech Emotion Recognition using MFCC and Hybrid Neural Networks. In International Joint Conference on Computational Intelligence (Vol. 1, pp. 366–373). Science and Technology Publications, Lda. https://doi.org/10.5220/0010707400003063
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.