Abstract:Adversarial training (AT) is a robust learning algorithm that can defend against adversarial attacks in the inference phase and mitigate the side effects of corrupted data in the training phase. As such, it has become an indispensable component of many artificial intelligence (AI) systems. However, in high-stake AI applications, it is crucial to understand AT's vulnerabilities to ensure reliable deployment. In this paper, we investigate AT's susceptibility to poisoning attacks, a type of malicious attack that manipulates training data to compromise the performance of the trained model. Previous work has focused on poisoning attacks against standard training, but little research has been done on their effectiveness against AT. To fill this gap, we design and test effective poisoning attacks against AT. Specifically, we investigate and design clean-label poisoning attacks, allowing attackers to imperceptibly modify a small fraction of training data to control the algorithm's behavior on a specific target data point. Additionally, we propose the clean-label untargeted attack, enabling attackers can attach tiny stickers on training data to degrade the algorithm's performance on all test data, where the stickers could serve as a signal against unauthorized data collection. Our experiments demonstrate that AT can still be poisoned, highlighting the need for caution when using vanilla AT algorithms in security-related applications. The code is at <a class="link-external link-https" href="https://github.com/zjfheart/Poison-adv-training.git" rel="external noopener nofollow">this https URL</a>.

UPTON: Preventing Authorship Leakage from Public Text Release via Data Poisoning

ALISON: Fast and Effective Stylometric Authorship Obfuscation

Concealed Data Poisoning Attacks on NLP Models

Keep It Private: Unsupervised Privatization of Online Text

Turning Generative Models Degenerate: The Power of Data Poisoning Attacks

JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models

Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws

Defending against Insertion-based Textual Backdoor Attacks via Attribution

UID as a Guiding Metric for Automated Authorship Obfuscation

Defending Against Authorship Identification Attacks

SHIELD: Thwarting Code Authorship Attribution

UTrace: Poisoning Forensics for Private Collaborative Learning

Assessing Vulnerabilities of Adversarial Learning Algorithm through Poisoning Attacks

Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models

TAROT: Task-Oriented Authorship Obfuscation Using Policy Optimization Methods

Text Laundering: Mitigating Malicious Features Through Knowledge Distillation of Large Foundation Models.

Avengers Ensemble! Improving Transferability of Authorship Obfuscation

A Girl Has A Name: Detecting Authorship Obfuscation

$A^{4}NT$: Author Attribute Anonymity by Adversarial Training of Neural Machine Translation

On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data Poisoning

Persistent Pre-Training Poisoning of LLMs