Abstract:Research of adversarial attacks is important for AI security because it shows the vulnerability of deep learning models and helps to build more robust models. Adversarial attacks on images are most widely studied, which include noise-based attacks, image editing-based attacks, and latent space-based attacks. However, the adversarial examples crafted by these methods often lack sufficient semantic information, making it challenging for humans to understand the failure modes of deep learning models under natural conditions. To address this limitation, we propose a natural language induced adversarial image attack method. The core idea is to leverage a text-to-image model to generate adversarial images given input prompts, which are maliciously constructed to lead to misclassification for a target model. To adopt commercial text-to-image models for synthesizing more natural adversarial images, we propose an adaptive genetic algorithm (GA) for optimizing discrete adversarial prompts without requiring gradients and an adaptive word space reduction method for improving query efficiency. We further used CLIP to maintain the semantic consistency of the generated images. In our experiments, we found that some high-frequency semantic information such as "foggy", "humid", "stretching", etc. can easily cause classifier errors. This adversarial semantic information exists not only in generated images but also in photos captured in the real world. We also found that some adversarial semantic information can be transferred to unknown classification tasks. Furthermore, our attack method can transfer to different text-to-image models (e.g., Midjourney, DALL-E 3, etc.) and image classifiers. Our code is available at: <a class="link-external link-https" href="https://github.com/zxp555/Natural-Language-Induced-Adversarial-Images" rel="external noopener nofollow">this https URL</a>.

Generating Fluent Adversarial Examples for Natural Languages.

Training NLI Models Through Universal Adversarial Attack

Generating Valid and Natural Adversarial Examples with Large Language Models

Generating Natural Language Adversarial Examples on a Large Scale with Generative Models

A Geometry-Inspired Attack for Generating Natural Language Adversarial Examples

Generating Fluent Chinese Adversarial Examples for Sentiment Classification

Frauds Bargain Attack: Generating Adversarial Text Samples via Word Manipulation Process

A Semantic, Syntactic, And Context-Aware Natural Language Adversarial Example Generator

Reversible Jump Attack to Textual Classifiers with Modification Reduction

Detecting Textual Adversarial Examples Based on Distributional Characteristics of Data Representations

JMA: a General Algorithm to Craft Nearly Optimal Targeted Adversarial Example

Learning to Attack: Towards Textual Adversarial Attacking in Real-world Situations

AdvExpander: Generating Natural Language Adversarial Examples by Expanding Text

Generating Adversarial Examples for Holding Robustness of Source Code Processing Models

Efficiently generating sentence-level textual adversarial examples with Seq2seq Stacked Auto-Encoder

Natural Language Induced Adversarial Images

SSCAE -- Semantic, Syntactic, and Context-aware natural language Adversarial Examples generator

Generating natural adversarial examples with universal perturbations for text classification

Towards Improving Adversarial Training of NLP Models

Enhancing Neural Models with Vulnerability Via Adversarial Attack.