Cat sounds classification with convolutional neural network

6Citations
Citations of this article
20Readers
Mendeley users who have this article in their library.

Abstract

In this study, we attempt to use a convolutional neural network (CNN) to identify cats’ different sounds. CNN is proven to classify different patterns from the spectro-temporal features of a sound and thus well suited for sound classification. We will perform data transformation using mel-frequency cepstral coefficients (MFCCs) to extract the sound frequency to apply this method. In MFCCs, each frequency bin is quasi-logarithmically spaced so that it resembles the resolution of the human auditory system compared to the spectrogram. We will be using four convolutional layers of CNN architecture with a pooling layer and dense layer as the output layer in our model. From the sound ontology Audio set, we can collect 595 different sound data classified into five categories of cat sounds, which we used to train our model. From our training process, our model can achieve a classification accuracy of 88.473254%. In the future, we look forward to improving our model accuracy by adding more data and even out each label to reduce overfitting. We would also like to implement a data augmentation method on our dataset to improve our model accuracy.

Cite

CITATION STYLE

APA

Ferdiana, R., Dicka, W. F., & Boediman, A. (2021). Cat sounds classification with convolutional neural network. International Journal on Electrical Engineering and Informatics, 13(3), 755–765. https://doi.org/10.15676/IJEEI.2021.13.3.15

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free