Abstract:This paper describes the USTC-KXDIGIT system submitted to the ASVspoof5 Challenge for Track 1 (speech deepfake detection) and Track 2 (spoofing-robust automatic speaker verification, SASV). Track 1 showcases a diverse range of technical qualities from potential processing algorithms and includes both open and closed conditions. For these conditions, our system consists of a cascade of a frontend feature extractor and a back-end classifier. We focus on extensive embedding engineering and enhancing the generalization of the back-end classifier model. Specifically, the embedding engineering is based on hand-crafted features and speech representations from a self-supervised model, used for closed and open conditions, respectively. To detect spoof attacks under various adversarial conditions, we trained multiple systems on an augmented training set. Additionally, we used voice conversion technology to synthesize fake audio from genuine audio in the training set to enrich the synthesis algorithms. To leverage the complementary information learned by different model architectures, we employed activation ensemble and fused scores from different systems to obtain the final decision score for spoof detection. During the evaluation phase, the proposed methods achieved 0.3948 minDCF and 14.33% EER in the close condition, and 0.0750 minDCF and 2.59% EER in the open condition, demonstrating the robustness of our submitted systems under adversarial conditions. In Track 2, we continued using the CM system from Track 1 and fused it with a CNN-based ASV system. This approach achieved 0.2814 min-aDCF in the closed condition and 0.0756 min-aDCF in the open condition, showcasing superior performance in the SASV system.

The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge

The huya multi-speaker and multi-style speech synthesis system for m2voc challenge 2020

The NPU-ASLP System for The ISCSLP 2022 Magichub Code-Swiching ASR Challenge

ICAGC 2024: Inspirational and Convincing Audio Generation Challenge 2024

UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge

The NPU System for the 2020 Personalized Voice Trigger Challenge

The NTU-AISG Text-to-speech System for Blizzard Challenge 2020

NWPU-ASLP System for the VoicePrivacy 2022 Challenge

KS-Net: Multi-band joint speech restoration and enhancement network for 2024 ICASSP SSI Challenge

The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024

The NPU-Elevoc Personalized Speech Enhancement System for ICASSP2023 DNS Challenge

The NPU-MSXF Speech-to-Speech Translation System for IWSLT 2023 Speech-to-Speech Translation Task

The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge

MSV Challenge 2022: NPU-HC Speaker Verification System for Low-resource Indian Languages

The SpeakIn System Description for CNSRC2022

Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge

USTC-KXDIGIT System Description for ASVspoof5 Challenge

Two-stage Neural Network for ICASSP 2023 Speech Signal Improvement Challenge

Overview of NIT HMM-based speech synthesis system for Blizzard Challenge 2009

TEA-PSE: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System for ICASSP 2022 DNS Challenge

DelightfulTTS: the Microsoft Speech Synthesis System for Blizzard Challenge 2021