Data augmentation and deep neural networks for the classification of Pakistani racial speakers recognition

Ammar Amjad; Lal Khan; Hsien Tsung Chang

Journal ArticleOPEN ACCESS

Data augmentation and deep neural networks for the classification of Pakistani racial speakers recognition

PeerJ Computer Science (2022) 8

DOI: 10.7717/PEERJ-CS.1053

4Citations

8Readers

Get full text

Abstract

Speech emotion recognition (SER) systems have evolved into an important method for recognizing a person in several applications, including e-commerce, everyday interactions, law enforcement, and forensics. The SER system’s efficiency depends on the length of the audio samples used for testing and training. However, the different suggested models successfully obtained relatively high accuracy in this study. Moreover, the degree of SER efficiency is not yet optimum due to the limited database, resulting in overfitting and skewing samples. Therefore, the proposed approach presents a data augmentation method that shifts the pitch, uses multiple window sizes, stretches the time, and adds white noise to the original audio. In addition, a deep model is further evaluated to generate a new paradigm for SER. The data augmentation approach increased the limited amount of data from the Pakistani racial speaker speech dataset in the proposed system. The seven-layer framework was employed to provide the most optimal performance in terms of accuracy compared to other multilayer approaches. The seven-layer method is used in existing works to achieve a very high level of accuracy. The suggested system achieved 97.32% accuracy with a 0.032% loss in the 75%:25% splitting ratio. In addition, more than 500 augmentation data samples were added. Therefore, the proposed approach results show that deep neural networks with data augmentation can enhance the SER performance on the Pakistani racial speech dataset.

Author supplied keywords

Cite

CITATION STYLE

APA

Amjad, A., Khan, L., & Chang, H. T. (2022). Data augmentation and deep neural networks for the classification of Pakistani racial speakers recognition. PeerJ Computer Science, 8. https://doi.org/10.7717/PEERJ-CS.1053

Data augmentation and deep neural networks for the classification of Pakistani racial speakers recognition

Abstract

Author supplied keywords

Cite

Register to see more suggestions