Abstract:We consider feature learning for efficient keyword spotting that can be applied in severely under-resourced settings. The objective is to support humanitarian relief programmes by the United Nations in parts of Africa in which almost no language resources are available. For rapid development in such languages, we rely on a small, easily-compiled set of isolated keywords. These keyword templates are applied to a large corpus of in-domain but untranscribed speech using dynamic time warping (DTW). The resulting DTW alignment scores are used to train a convolutional neural network (CNN) which is orders of magnitude more computationally efficient and suitable for real-time application. We optimise this neural network keyword spotter by identifying robust acoustic features in this almost zero-resource setting. First, we incorporate information from well-resourced but unrelated languages using a multilingual bottleneck feature (BNF) extractor. Next, we consider features extracted from an autoencoder (AE) trained on in-domain but untranscribed data. Finally, we consider correspondence autoencoder (CAE) features which are fine-tuned on the small set of in-domain labelled data. Experiments in South African English and Luganda, a low-resource language, show that BNF and CAE features achieve a 5% relative performance improvement over baseline MFCCs. However, using BNFs as input to the CAE results in a more than 27% relative improvement over MFCCs in ROC area-under-the-curve (AUC) and more than twice as many top-10 retrievals. We show that, using these features, the CNN-DTW keyword spotter performs almost as well as the DTW keyword spotter while outperforming a baseline CNN trained only on the keyword templates. The CNN-DTW keyword spotter using BNF-derived CAE features represents an efficient approach with competitive performance suited to rapid deployment in a severely under-resourced scenario.

Multilingual bottleneck features for subword modeling in zero-resource languages

Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced Languages

Improved acoustic word embeddings for zero-resource languages using multilingual transfer

Adversarial Multilingual Training for Low-Resource Speech Recognition.

Multilingual acoustic word embedding models for processing zero-resource languages

ASR-free CNN-DTW keyword spotting using multilingual bottleneck features for almost zero-resource languages

Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations

Universal Cross-Lingual Data Generation for Low Resource ASR

Finding Better Subword Segmentation for Neural Machine Translation

Towards Lingua Franca Named Entity Recognition with BERT

Towards Fully Bilingual Deep Language Modeling

Towards Building ASR Systems for the Next Billion Users

Low-resource Bilingual Dialect Lexicon Induction with Large Language Models

Feature learning for efficient ASR-free keyword spotting in low-resource languages

Multilingual Speech Translation with Efficient Finetuning of Pretrained Models

Bilingual Lexicon Induction from Non-Parallel Data with Minimal Supervision.

Language-invariant Bottleneck Features from Adversarial End-to-end Acoustic Models for Low Resource Speech Recognition.

Establishing degrees of closeness between audio recordings along different dimensions using large-scale cross-lingual models

A General Procedure for Improving Language Models in Low-Resource Speech Recognition

A Balanced Data Approach for Evaluating Cross-Lingual Transfer: Mapping the Linguistic Blood Bank

Leveraging Multilingual Self-Supervised Pretrained Models for Sequence-to-Sequence End-to-End Spoken Language Understanding