Abstract
Convolutional Neural Networks (CNNs) have achieved tremendous success in image classification tasks. However, CNNs are vulnerable to adversarial attacks, such as applying imperceptible perturbations on the legitimate images. To address the security threats posed by these adversarial attacks, many defense techniques have been proposed. Adversarial training has been shown to be effective in enhancing CNNs robustness against adversarial samples. However, the trade-off between robustness and classification accuracy in adversarial training cannot be overlooked. In this paper, we propose a novel approach to adversarial training that simultaneously trains the model using both real images and minimally perturbed borderline adversaries. These borderline adversaries were generated during the training process, using the shortest successful perturbations for each individual training sample at specific training states. Instead of training with a fixed ϵ that applies uniform perturbations to all training samples, this shortest successful perturbation is adaptive to the network’s training state and is automatically determined. The rationale behind this approach is that the decision boundary will be less distorted by these additional adversaries, which helps maintain the classification accuracy while improving adversarial robustness. Preliminary experiments conducted on Cholec80 dataset for surgical tool recognition showed that this method achieved 2–7% improvement in both adversarial robustness and accuracy compared to other adversarial training methods, while also reducing overlap in the classification regions after adversarial training.
Author supplied keywords
Cite
CITATION STYLE
Ding, N., & Möller, K. (2025). Adversarial training with borderline samples. Journal of Supercomputing, 81(8). https://doi.org/10.1007/s11227-025-07477-3
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.