Inspecting Prediction Confidence for Detecting Black-Box Backdoor Attacks

16Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Backdoor attacks have been shown to be a serious security threat against deep learning models, and various defenses have been proposed to detect whether a model is backdoored or not. However, as indicated by a recent black-box attack, existing defenses can be easily bypassed by implanting the backdoor in the frequency domain. To this end, we propose a new defense DTINSPECTOR against black-box backdoor attacks, based on a new observation related to the prediction confidence of learning models. That is, to achieve a high attack success rate with a small amount of poisoned data, backdoor attacks usually render a model exhibiting statistically higher prediction confidences on the poisoned samples. We provide both theoretical and empirical evidence for the generality of this observation. DTINSPECTOR then carefully examines the prediction confidences of data samples, and decides the existence of backdoor using the shortcut nature of backdoor triggers. Extensive evaluations on six backdoor attacks, four datasets, and three advanced attacking types demonstrate the effectiveness of the proposed defense.

Cite

CITATION STYLE

APA

Wang, T., Yao, Y., Xu, F., Xu, M., An, S., & Wang, T. (2024). Inspecting Prediction Confidence for Detecting Black-Box Backdoor Attacks. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, pp. 274–282). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v38i1.27780

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free