Abstract:Establishing trust and helping experts debug and understand the inner workings of deep learning models, interpretation methods are increasingly coupled with these models, building interpretable deep learning systems. However, adversarial attacks pose a significant threat to public trust by making interpretations of deep learning models confusing and difficult to understand. In this paper, we present a novel Single-class target-specific ADVersarial attack called SingleADV. The goal of SingleADV is to generate a universal perturbation that deceives the target model into confusing a specific category of objects with a target category while ensuring highly relevant and accurate interpretations. The universal perturbation is stochastically and iteratively optimized by minimizing the adversarial loss that is designed to consider both the classifier and interpreter costs in targeted and non-targeted categories. In this optimization framework, ruled by the first- and second-moment estimations, the desired loss surface promotes high confidence and interpretation scores of adversarial samples. By avoiding unintended misclassification of samples from other categories, SingleADV enables more effective targeted attacks on interpretable deep learning systems in both white-box and black-box scenarios. To evaluate the effectiveness of SingleADV, we conduct experiments using four different model architectures (ResNet-50, VGG-16, DenseNet-169, and Inception-V3) coupled with three interpretation models (CAM, Grad, and MASK). Through extensive empirical evaluation, we demonstrate that SingleADV effectively deceives target deep learning models and their associated interpreters under various conditions and settings. Our results show that the performance of SingleADV is effective, with an average attack success rate of 74% and prediction confidence exceeding 77% on successful adversarial samples. Furthermore, we discuss several countermeasures against SingleADV, including a transfer-based learning approach and existing preprocessing defenses.

Interpretable adversarial example detection via high-level concept activation vector

Fooling Neural Network Interpretations - Adversarial Noise to Attack Images.

Protego: Detecting Adversarial Examples for Vision Transformers Via Intrinsic Capabilities

Adversarial Attacks Hidden in Plain Sight

Interpreting Adversarial Examples by Activation Promotion and Suppression

Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples

Adversarial Examples Detection Beyond Image Space.

Demiguise Attack: Crafting Invisible Semantic Adversarial Perturbations with Perceptual Similarity

Model-agnostic Adversarial Example Detection via High-Frequency Amplification

Imperceptible Adversarial Attack via Invertible Neural Networks

Adversarial example detection based on saliency map features

Hlr: Generating Adversarial Examples By High-Level Representations

Attacks Meet Interpretability: Attribute-steered Detection of Adversarial Samples

The Anatomy of Adversarial Attacks: Concept-based XAI Dissection

SingleADV: Single-Class Target-Specific Attack Against Interpretable Deep Learning Systems

Towards Imperceptible and Robust Adversarial Example Attacks Against Neural Networks

Proper Network Interpretability Helps Adversarial Robustness in Classification

Searching for the Essence of Adversarial Perturbations

New Adversarial Image Detection Based on Sentiment Analysis

Brain-inspired reverse adversarial examples.

Interpreting Adversarial Examples in Deep Learning: A Review