Abstract:Adversarial examples to speaker recognition (SR) systems are generated by adding a carefully crafted noise to the speech signal to make the system fail while being imperceptible to humans. Such attacks pose severe security risks, making it vital to deep-dive and understand how much the state-of-the-art SR systems are vulnerable to these attacks. Moreover, it is of greater importance to propose defenses that can protect the systems against these attacks. Addressing these concerns, this paper at first investigates how state-of-the-art x-vector based SR systems are affected by white-box adversarial attacks, i.e., when the adversary has full knowledge of the system. x-Vector based SR systems are evaluated against white-box adversarial attacks common in the literature like fast gradient sign method (FGSM), basic iterative method (BIM)--a.k.a. iterative-FGSM--, projected gradient descent (PGD), and Carlini-Wagner (CW) attack. To mitigate against these attacks, the paper proposes four pre-processing defenses. It evaluates them against powerful adaptive white-box adversarial attacks, i.e., when the adversary has full knowledge of the system, including the defense. The four pre-processing defenses--viz. randomized smoothing, DefenseGAN, variational autoencoder (VAE), and Parallel WaveGAN vocoder (PWG) are compared against the baseline defense of adversarial training. Conclusions indicate that SR systems were extremely vulnerable under BIM, PGD, and CW attacks. Among the proposed pre-processing defenses, PWG combined with randomized smoothing offers the most protection against the attacks, with accuracy averaging 93% compared to 52% in the undefended system and an absolute improvement >90% for BIM attacks with $L_\infty>0.001$ and CW attack.

An attack-agnostic defense method against adversarial attacks on speaker verification by fusing downsampling and upsampling of speech signals

Echo: Reverberation-based Fast Black-Box Adversarial Attacks on Intelligent Audio Systems.

FenceSitter: Black-box, Content-Agnostic, and Synchronization-Free Enrollment-Phase Attacks on Speaker Recognition Systems

UltraBD: Backdoor Attack against Automatic Speaker Verification Systems via Adversarial Ultrasound

Spoofing Speaker Verification System by Adversarial Examples Leveraging the Generalized Speaker Difference.

Defending Adversarial Attacks on Cloud-aided Automatic Speech Recognition Systems.

BypTalker: an Adaptive Adversarial Example Attack to Bypass Prefilter-enabled Speaker Recognition

Defending Against Adversarial Attacks in Speaker Verification Systems

Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning

Query-Efficient Adversarial Attack with Low Perturbation Against End-to-End Speech Recognition Systems

Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems

Towards Understanding and Mitigating Audio Adversarial Examples for Speaker Recognition

Adversarial Sample Detection for Speaker Verification by Neural Vocoders

Adversarial defense for deep speaker recognition using hybrid adversarial training

Adversarial Attack and Defense Strategies of Speaker Recognition Systems: A Survey

Push the Limit of Adversarial Example Attack on Speaker Recognition in Physical Domain

Enrollment-stage Backdoor Attacks on Speaker Recognition Systems via Adversarial Ultrasound

Diffusion-Based Adversarial Purification for Speaker Verification

A Non-intrusive and Adaptive Speaker De-Identification Scheme Using Adversarial Examples

The defender's perspective on automatic speaker verification: An overview

Scalable Ensemble-based Detection Method against Adversarial Attacks for speaker verification