Abstract:Research of adversarial attacks is important for AI security because it shows the vulnerability of deep learning models and helps to build more robust models. Adversarial attacks on images are most widely studied, which include noise-based attacks, image editing-based attacks, and latent space-based attacks. However, the adversarial examples crafted by these methods often lack sufficient semantic information, making it challenging for humans to understand the failure modes of deep learning models under natural conditions. To address this limitation, we propose a natural language induced adversarial image attack method. The core idea is to leverage a text-to-image model to generate adversarial images given input prompts, which are maliciously constructed to lead to misclassification for a target model. To adopt commercial text-to-image models for synthesizing more natural adversarial images, we propose an adaptive genetic algorithm (GA) for optimizing discrete adversarial prompts without requiring gradients and an adaptive word space reduction method for improving query efficiency. We further used CLIP to maintain the semantic consistency of the generated images. In our experiments, we found that some high-frequency semantic information such as "foggy", "humid", "stretching", etc. can easily cause classifier errors. This adversarial semantic information exists not only in generated images but also in photos captured in the real world. We also found that some adversarial semantic information can be transferred to unknown classification tasks. Furthermore, our attack method can transfer to different text-to-image models (e.g., Midjourney, DALL-E 3, etc.) and image classifiers. Our code is available at: <a class="link-external link-https" href="https://github.com/zxp555/Natural-Language-Induced-Adversarial-Images" rel="external noopener nofollow">this https URL</a>.

Distilling Adversarial Prompts from Safety Benchmarks: Report for the Adversarial Nibbler Challenge

Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation

SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution

SneakyPrompt: Jailbreaking Text-to-image Generative Models

Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts

On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts

Black Box Adversarial Prompting for Foundation Models

Divide-and-Conquer Attack: Harnessing the Power of LLM to Bypass the Censorship of Text-to-Image Generation Model

When Image Generation Goes Wrong: A Safety Analysis of Stable Diffusion Models

Backdooring Bias into Text-to-Image Models

BSPA: Exploring Black-box Stealthy Prompt Attacks Against Image Generators

Natural Language Induced Adversarial Images

Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion models

Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding

ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users

Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization

Universal Prompt Optimizer for Safe Text-to-Image Generation

A Prompt-Based Approach to Adversarial Example Generation and Robustness Enhancement

Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation