Abstract:Research of adversarial attacks is important for AI security because it shows the vulnerability of deep learning models and helps to build more robust models. Adversarial attacks on images are most widely studied, which include noise-based attacks, image editing-based attacks, and latent space-based attacks. However, the adversarial examples crafted by these methods often lack sufficient semantic information, making it challenging for humans to understand the failure modes of deep learning models under natural conditions. To address this limitation, we propose a natural language induced adversarial image attack method. The core idea is to leverage a text-to-image model to generate adversarial images given input prompts, which are maliciously constructed to lead to misclassification for a target model. To adopt commercial text-to-image models for synthesizing more natural adversarial images, we propose an adaptive genetic algorithm (GA) for optimizing discrete adversarial prompts without requiring gradients and an adaptive word space reduction method for improving query efficiency. We further used CLIP to maintain the semantic consistency of the generated images. In our experiments, we found that some high-frequency semantic information such as "foggy", "humid", "stretching", etc. can easily cause classifier errors. This adversarial semantic information exists not only in generated images but also in photos captured in the real world. We also found that some adversarial semantic information can be transferred to unknown classification tasks. Furthermore, our attack method can transfer to different text-to-image models (e.g., Midjourney, DALL-E 3, etc.) and image classifiers. Our code is available at: <a class="link-external link-https" href="https://github.com/zxp555/Natural-Language-Induced-Adversarial-Images" rel="external noopener nofollow">this https URL</a>.

Language Model Agnostic Gray-Box Adversarial Attack on Image Captioning

Fooled by imagination: adversarial attack to image captioning via perturbation in complex domain

Restricted-Area Adversarial Example Attack for Image Captioning Model

AICAttack: Adversarial Image Captioning Attack with Attention-Based Optimization

Exact Adversarial Attack to Image Captioning via Structured Output Learning with Latent Variables

Mimic and Fool: A Task-Agnostic Adversarial Attack

Natural Language Induced Adversarial Images

Generating Natural Language Adversarial Examples on a Large Scale with Generative Models

Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks

Fooling Vision and Language Models Despite Localization and Attention Mechanism

Undermining Image and Text Classification Algorithms Using Adversarial Attacks

Learning Transferable Perturbations for Image Captioning

AnyAttack: Targeted Adversarial Attacks on Vision-Language Models toward Any Images

Mutual-modality Adversarial Attack with Semantic Perturbation

Attack as the Best Defense: Nullifying Image-to-image Translation GANs via Limit-aware Adversarial Attack

Stealthy Targeted Backdoor Attacks against Image Captioning

Cascade & allocate: A cross-structure adversarial attack against models fusing vision and language

Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training

Object-oriented backdoor attack against image captioning

Ask, Attend, Attack: A Effective Decision-Based Black-Box Targeted Attack for Image-to-Text Models