Abstract
Deep neural networks (DNNs) for image classification remain vulnerable to adversarial perturbations–subtle input manipulations that induce catastrophic misclassifications. To address this issue, we propose the Adversarial Image Rectifier (AIR), a linguistically inspired detection and mitigation framework that enhances DNN robustness by intercepting and inverting adversarial perturbations at the feature level. Unlike existing defenses, AIR operates without prior knowledge of attack patterns: it first encodes hierarchical hidden-layer feature maps of a DNN into semantically structured sentence representations, then identifies adversarial inputs through “sentiment” anomalies in these sentences–a linguistic metaphor for subtle adversarial traces. Crucially, we pinpoint a pivotal intermediate layer where adversarial perturbations dominantly propagate and train a lightweight rectifier network to selectively nullify adversarial features at this layer while preserving benign semantics. Extensive experiments on Tiny-ImageNet, CIFAR-10, SVHN, and MS COCO demonstrate that AIR achieves a correction rate of up to 95.02% and 94.62% when defending against known attacks and unknown attacks, respectively, significantly surpassing existing defense techniques.
Author supplied keywords
Cite
CITATION STYLE
Wang, Y., Song, J., Li, T., Xin, Y., Li, H., & Wei, N. (2025). Rectifying Multi-Attack Adversarial Perturbations in Deep Neural Network based Image Classifier. ACM Transactions on Privacy and Security, 28(4). https://doi.org/10.1145/3765757
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.