Speech Emotion Recognition using MFCC and Hybrid Neural Networks

1Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Speech emotion recognition is a challenging task and feature extraction plays an important role in effectively classifying speech into different emotions. In this paper, we apply traditional feature extraction methods like MFCC for feature extraction from audio files. Instead of using traditional machine learning approaches like SVM to classify audio files, we investigate different neural network architectures. Our baseline model implemented as a convolutional neural network results in 60% classification accuracy. We propose a hybrid neural network architecture based on Convolutional and Long Short-Term Memory (ConvLSTM) networks to capture spatial and sequential information of audio files. Our experimental results show that our ComvLSTM model has achieved an accuracy of 59%. We improved our model with data augmentation techniques and re-trained it with augmented dataset. The classification accuracy achieves 91% for multi-class classification of RAVDESS dataset outperforming the accuracy of state-of-the-art multi-class classification models that used the similar data.

Cite

CITATION STYLE

APA

Badr, Y., Mukherjee, P., & Thumati, S. M. (2021). Speech Emotion Recognition using MFCC and Hybrid Neural Networks. In International Joint Conference on Computational Intelligence (Vol. 1, pp. 366–373). Science and Technology Publications, Lda. https://doi.org/10.5220/0010707400003063

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free