Abstract:This paper describes the USTC-KXDIGIT system submitted to the ASVspoof5 Challenge for Track 1 (speech deepfake detection) and Track 2 (spoofing-robust automatic speaker verification, SASV). Track 1 showcases a diverse range of technical qualities from potential processing algorithms and includes both open and closed conditions. For these conditions, our system consists of a cascade of a frontend feature extractor and a back-end classifier. We focus on extensive embedding engineering and enhancing the generalization of the back-end classifier model. Specifically, the embedding engineering is based on hand-crafted features and speech representations from a self-supervised model, used for closed and open conditions, respectively. To detect spoof attacks under various adversarial conditions, we trained multiple systems on an augmented training set. Additionally, we used voice conversion technology to synthesize fake audio from genuine audio in the training set to enrich the synthesis algorithms. To leverage the complementary information learned by different model architectures, we employed activation ensemble and fused scores from different systems to obtain the final decision score for spoof detection. During the evaluation phase, the proposed methods achieved 0.3948 minDCF and 14.33% EER in the close condition, and 0.0750 minDCF and 2.59% EER in the open condition, demonstrating the robustness of our submitted systems under adversarial conditions. In Track 2, we continued using the CM system from Track 1 and fused it with a CNN-based ASV system. This approach achieved 0.2814 min-aDCF in the closed condition and 0.0756 min-aDCF in the open condition, showcasing superior performance in the SASV system.

The XMU System for Audio-Visual Diarization and Recognition in MISP Challenge 2022.

The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition

Summary on the Multimodal Information Based Speech Processing (MISP) 2022 Challenge.

The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge

Summary on the Multimodal Information-Based Speech Processing (MISP) 2023 Challenge

Cross-modal Mask Fusion and Modality-Balanced Audio-Visual Speech Recognition

The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And Results

The xmuspeech system for multi-channel multi-party meeting transcription challenge

The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction

XMUSPEECH System for VoxCeleb Speaker Recognition Challenge 2021

The DKU Post-Challenge Audio-Visual Wake Word Spotting System for the 2021 MISP Challenge: Deep Analysis

The FlySpeech Audio-Visual Speaker Diarization System for MISP Challenge 2022

USTC-KXDIGIT System Description for ASVspoof5 Challenge

Challenges and Insights: Exploring 3D Spatial Features and Complex Networks on the MISP Dataset

The SpeakIn System Description for CNSRC2022

An Investigation into Audio–Visual Speech Recognition under a Realistic Home–TV Scenario

The CUHK-TENCENT speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge

An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge

Royalflush Speaker Diarization System for ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge

SIR-Progressive Audio-Visual TF-Gridnet with ASR-Aware Selector for Target Speaker Extraction in MISP 2023 Challenge