Model Stealing Defense against Exploiting Information Leak through the Interpretation of Deep Neural Nets

13Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.

Abstract

Model stealing techniques allow adversaries to create attack models that mimic the functionality of black-box machine learning models, querying only class membership or probability outcomes. Recently, interpretable AI is getting increasing attention, to enhance our understanding of AI models, provide additional information for diagnoses, or satisfy legal requirements. However, it has been recently reported that providing such additional information can make AI models more vulnerable to model stealing attacks. In this paper, we propose DeepDefense, the first defense mechanism that protects an AI model against model stealing attackers exploiting both class probabilities and interpretations. DeepDefense uses a misdirection model to hide the critical information of the original model against model stealing attacks, with minimal degradation on both the class probability and the interpretability of prediction output. DeepDefense is highly applicable for any model stealing scenario since it makes minimal assumptions about the model stealing adversary. In our experiments, DeepDefense shows significantly higher defense performance than the existing state-of-the-art defenses on various datasets and interpreters.

Cite

CITATION STYLE

APA

Lee, J., Han, S., & Lee, S. (2022). Model Stealing Defense against Exploiting Information Leak through the Interpretation of Deep Neural Nets. In IJCAI International Joint Conference on Artificial Intelligence (pp. 710–716). International Joint Conferences on Artificial Intelligence. https://doi.org/10.24963/ijcai.2022/100

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free