Speech-based emotion recognition and speaker identification: Static vs. dynamic mode of speech representation

7Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

In this paper we present the performance of different machine learning algorithms for the problems of speech-based Emotion Recognition (ER) and Speaker Identification (SI) in static and dynamic modes of speech signal representation. We have used a multi-corporal, multi-language approach in the study. 3 databases for the problem of SI and 4 databases for the ER task of 3 different languages (German, English and Japanese) have been used in our study to evaluate the models. More than 45 machine learning algorithms were applied to these tasks in both modes and the results alongside discussion are presented here.

Cite

CITATION STYLE

APA

Sidorov, M., Minker, W., & Semenkin, E. S. (2016). Speech-based emotion recognition and speaker identification: Static vs. dynamic mode of speech representation. Journal of Siberian Federal University - Mathematics and Physics, 9(4), 518–523. https://doi.org/10.17516/1997-1397-2016-9-4-518-523

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free