Abstract
Existing work finds it challenging for adversarial examples to transfer among different synthetic speech detectors because of cross-feature and cross-model. To enhance the transferability of adversarial examples, we propose a spectral saliency analysis method and gain insight into the underlying detection mechanisms of existing detectors for the first time. These insights offer an interpretable basis for why adversarial examples are challenging to transfer between synthetic speech detection models. Then we further propose a two-stage adversarial attack framework. Specifically, the first stage leverages insights into the model detection mechanism to design a random time-frequency masking module, the random offset module, and 1D convolution to generate transferable and robust adversarial examples. In the second stage, to mitigate the problem of obvious noise in the low-energy frames of the carrier in existing adversarial attacks, we perform secondary optimization on frames below the Signal-Noise-Rate threshold to enhance its auditory quality. Extensive experimental results demonstrate that the proposed method significantly enhances the transferability and robustness of adversarial examples, while simultaneously preserving the acoustic quality compared to typical approaches.
Author supplied keywords
Cite
CITATION STYLE
Deng, J., Ye, D., Li, J., Liu, Z., Tang, L., & Zhang, Y. (2025). The Interpretable and Transferable Adversarial Attack against Synthetic Speech Detectors. ACM Transactions on Multimedia Computing, Communications and Applications, 21(5). https://doi.org/10.1145/3727341
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.