Cosine Scoring with Uncertainty for Neural Speaker Embedding

Qiongqiong Wang,Kong Aik Lee
DOI: https://doi.org/10.1109/LSP.2024.3375080
2024-03-11
Abstract:Uncertainty modeling in speaker representation aims to learn the variability present in speech utterances. While the conventional cosine-scoring is computationally efficient and prevalent in speaker recognition, it lacks the capability to handle uncertainty. To address this challenge, this paper proposes an approach for estimating uncertainty at the speaker embedding front-end and propagating it to the cosine scoring back-end. Experiments conducted on the VoxCeleb and SITW datasets confirmed the efficacy of the proposed method in handling uncertainty arising from embedding estimation. It achieved improvement with 8.5% and 9.8% average reductions in EER and minDCF compared to the conventional cosine similarity. It is also computationally efficient in practice.
Sound,Machine Learning,Audio and Speech Processing
What problem does this paper attempt to address?
This paper focuses on the handling of uncertainty in speaker embedding in speech recognition. The traditional cosine similarity calculation has limitations in dealing with this uncertainty. The paper proposes a new approach that propagates uncertainty from the front-end embedding estimation to the back-end cosine similarity scoring process. Experimental results on the VoxCeleb and SITW datasets demonstrate the effectiveness of this approach in handling uncertainty caused by embedding estimation, with an average reduction of 8.5% in equal error rate (EER) and 9.8% in minimum detection cost (minDCF), improving the performance of speaker recognition. Furthermore, this approach also has high computational efficiency in practical applications.