Acoustic Word Embedding System for Code-Switching Query-by-example Spoken Term Detection

Murong Ma,Haiwei Wu,Xuyang Wang,Lin Yang,Junjie Wang,Ming Li
DOI: https://doi.org/10.1109/iscslp49672.2021.9362056
2021-01-24
Abstract:In this paper, we propose a deep convolutional neural network-based acoustic word embedding system for code-switching query by example spoken term detection. Different from previous configurations, we combine audio data in two languages for training instead of only using one single language. We trans-form the acoustic features of keyword templates and searching content segments obtained in a sliding manner to fixed-dimensional vectors and calculate the distances between them. An auxiliary variability-invariant loss is also applied to training data within the same word but different speakers. This strategy is used to prevent the extractor from encoding undesired speaker- or accent-related information into the acoustic word embeddings. Experimental results show that our proposed sys-tem produces promising searching results in the code-switching test scenario. With the employment of variability-invariant loss, the searching performance is further enhanced.
What problem does this paper attempt to address?