Abstract:Transferability of adversarial examples is critical for black-box deep learning model attacks. While most existing studies focus on enhancing the transferability of untargeted adversarial attacks, few of them studied how to generate transferable targeted adversarial examples that can mislead models into predicting a specific class. Moreover, existing transferable targeted adversarial attacks usually fail to sufficiently characterize the target class distribution, thus suffering from limited transferability. In this paper, we propose the Transferable Targeted Adversarial Attack (TTAA), which can capture the distribution information of the target class from both label-wise and feature-wise perspectives, to generate highly transferable targeted adversarial examples. To this end, we design a generative adversarial training framework consisting of a generator to produce targeted adversarial examples, and feature-label dual discriminators to distinguish the generated adversarial examples from the target class images. Specifically, we design the label discriminator to guide the adversarial examples to learn label-related distribution information about the target class. Meanwhile, we design a feature discriminator, which extracts the feature-wise information with strong cross-model consistency, to enable the adversarial examples to learn the transferable distribution information. Furthermore, we introduce the random perturbation dropping to further enhance the transferability by augmenting the diversity of adversarial examples used in the training process. Experiments demonstrate that our method achieves excellent performance on the transferability of targeted adversarial examples. The targeted fooling rate reaches 95.13% when transferred from VGG-19 to DenseNet-121, which significantly outperforms the state-of-the-art methods.

Transferable Adversarial Examples Can Efficiently Fool Topic Models

Towards Efficient Data Free Blackbox Adversarial Attack

Misleading Sentiment Analysis: Generating Adversarial Texts by the Ensemble Word Addition Algorithm

Understanding Model Ensemble in Transferable Adversarial Attack

Understanding and Enhancing the Transferability of Adversarial Examples

On the Transferability of Adversarial Attacksagainst Neural Text Classifier

Rethinking Model Ensemble in Transfer-based Adversarial Attacks

Generating Universal Language Adversarial Examples by Understanding and Enhancing the Transferability Across Neural Models

Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples

Towards Transferable Targeted Adversarial Examples

Bag of Tricks to Boost Adversarial Transferability

Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks

Boosting the Transferability of Ensemble Adversarial Attack via Stochastic Average Variance Descent

Delving into Transferable Adversarial Examples and Black-box Attacks

Transfer Adversarial Attacks Across Industrial Intelligent Systems.

Towards Transferable Targeted Attack.

How to choose your best allies for a transferable attack?

GCSA: A New Adversarial Example-Generating Scheme Towards Black-Box Adversarial Attacks

Enhancing Adversarial Transferability with Adversarial Weight Tuning

An Adaptive Model Ensemble Adversarial Attack for Boosting Adversarial Transferability

Enhancing Adversarial Attacks: The Similar Target Method