Abstract
Audio deepfakes, a subset of deepfake technology, employ machine learning or deep learning to create deceptive audio content by synthesizing authentic recordings. Such deepfakes not only fosters the dissemination of misinformation but also empowers identity theft while compromising individual privacy. Discerning between counterfeit and authentic audio content poses escalating challenges for digital forensic analysts. The proposed paper develops a robust deep learning model that harnesses fusion approach with a spectrum of diverse audio spectral features to effectively detect deepfake audios. By employing a fusion strategy, the developed model ensembles predictions from two pre-trained networks, CIFAR-10 and ResNet50. Additionally, it capitalizes on a diverse array of spectral audio features- Mel Frequency Cepstral Coefficients (MFCC), Constant-Q Cepstral Coefficients (CQCC), Mel-Spectrogram, and Spectral Centroid, for extraction of crucial details from raw audio data. Accuracy of proposed model is assessed on more recent and widely used FoR dataset having three sub-datasets of 195,000 audio samples. Experimental results reveal that proposed model achieves superior performance, boasting an accuracy of 99.12%, precision of 97.54%, recall of 98.44%, and F1 score of 98.12% when utilizing MFCC feature as compared to other audio features. Moreover, the model undergoes an accuracy-centric quantitative assessment, surpassing eight state-of-the-art audio detection models, including DNN, DeepSonar, STN, TCN, SVM, CNN, KNN, and RF.
Cite
CITATION STYLE
Kaur, N., Dixit, A., & Kingra, S. (2025). A Deep Learning Fusion Model Leveraging Spectral Features for Audio Deepfake Detection. International Journal of Advanced Networking and Applications, 16(05), 6551–6560. https://doi.org/10.35444/ijana.2025.16503
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.