Frequency domain-based detection of generated audio

22Citations
Citations of this article
33Readers
Mendeley users who have this article in their library.

Abstract

Attackers may manipulate audio with the intent of presenting falsified reports, changing an opinion of a public figure, and winning influence and power. The prevalence of inauthentic multimedia continues to rise, so it is imperative to develop a set of tools that determines the legitimacy of media. We present a method that analyzes audio signals to determine whether they contain real human voices or fake human voices (i.e., voices generated by neural acoustic and waveform models). Instead of analyzing the audio signals directly, the proposed approach converts the audio signals into spectrogram images displaying frequency, intensity, and temporal content and evaluates them with a Convolutional Neural Network (CNN). Trained on both genuine human voice signals and synthesized voice signals, we show our approach achieves high accuracy on this classification task.

Cite

CITATION STYLE

APA

Bartusiak, E. R., & Delp, E. J. (2021). Frequency domain-based detection of generated audio. In IS and T International Symposium on Electronic Imaging Science and Technology (Vol. 2021). Society for Imaging Science and Technology. https://doi.org/10.2352/ISSN.2470-1173.2021.4.MWSF-273

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free