Abstract:Recently, it was found that deep neural networks (DNNs) are susceptible to adversarial input perturbations. Most defense strategies adopt the denoising method based on preprocessing, which mitigates the impacts of adversarial perturbations on DNNs by learning the distributions of nonadversarial datasets and projecting adversarial inputs into the learned nonadversarial manifolds. However, existing defense strategies commonly focus on reconstructing clean images while ignoring the role of adversarial perturbations, which results in the reconstructed images failing to achieve the visual quality and classification accuracy of the original clean images, and the induced adversarial robustness improvement is limited. This paper proposes a feature decoupling-interaction network (FDIN), which introduces the concepts of clean features and adversarial features to separate the two kinds of features from the input adversarial examples (AEs) in a feature decoupling-interaction manner. The clean features are used to reconstruct the input image so that it is infinitely close to the original clean image, and the adversarial features are used to reconstruct the adversarial perturbations. Adversarial perturbations are removed from the adversarial examples across multiple cross cycles to improve further the reconstructed image's visual quality and classification accuracy. The features of the original clean image are used as prior knowledge to guide the network to learn the clean features of the adversarial examples and improve the classification accuracy of the model on the clean examples. In addition, a classification loss function based on the Carlini & Wagner (CW) attack algorithm is used instead of the conventional cross-entropy loss function to improve the adversarial robustness of the FDIN. The experimental results show that the proposed method achieves better defense performance than the current state-of-the-art methods on both standard tests and various attack tests and even exceeds the test accuracy of the target classifier on the original test set.

Removing Adversarial Noise in Class Activation Feature Space

Adversarial perturbation denoising utilizing common characteristics in deep feature space

Feature Denoising for Improving Adversarial Robustness

Analyzing the Noise Robustness of Deep Neural Networks

A Universal Defense Strategy Against Adversarial Attacks Based on Attention-Guided

Enhance DNN Adversarial Robustness and Efficiency via Injecting Noise to Non-Essential Neurons

Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals

Procedural Noise Adversarial Examples for Black-Box Attacks on Deep Convolutional Networks

LDN-RC: a Lightweight Denoising Network with Residual Connection to Improve Adversarial Robustness

Feature decoupling and interaction network for defending against adversarial examples

Training Robust Deep Neural Networks via Adversarial Noise Propagation

Efficiently Finding Adversarial Examples with DNN Preprocessing

Preprocessing-based Adversarial Defense for Object Detection Via Feature Filtration.

A Noise-Sensitivity-Analysis-Based Test Prioritization Technique for Deep Neural Networks

An Efficient Pre-processing Method to Eliminate Adversarial Effects

General Adversarial Defense Against Black-box Attacks Via Pixel Level and Feature Level Distribution Alignments

Generating Adversarial Attacks in the Latent Space

Propagated Perturbation of Adversarial Attack for well-known CNNs: Empirical Study and its Explanation

Detecting adversarial samples by noise injection and denoising

Enhancing Robust Representation in Adversarial Training: Alignment and Exclusion Criteria

Are You Confident That You Have Successfully Generated Adversarial Examples?