Abstract:Information retrieval from spoken audio has attracted the attention of a number of research groups, in part driven by the recent NIST Spoken Term Detection (STD) evaluation. A common approach is to split the task into two stages. In the first, a large vocabulary continuous speech recognition (LVCSR) system is used to generate a word or phone lattice corresponding to the audio, and in the second, lattice search is used to determine likely occurrences of the search terms. Searching a word-based lattice works well for terms which occur in the LVCSR system's vocabulary. However, search terms naturally have a tendency toward proper nouns, which leads to higher out-of-vocabulary (OOV) rates than found in transcription tasks. A standard method for dealing with OOV terms is to generate a phone sequence corresponding to the terms, which may be then be searched for in a phone lattice. In this work, we propose using context-dependent graphemes (CDG) as sub-word units for spoken term detection, in particular for out-of-vocabulary search terms. In essence, this approach moves pronunciation modelling away from the letter-to-sound rules which are used to generate phone strings, and into the Gaus-sian mixture models which describe the observation space. This removes the need to make potentially error-prone hard decisions at an early stage of processing. In addition, words which have multiple pronunciations have a single grapheme representation which simplifies the subsequent search. Large text corpora can be used to train long-span grapheme-based language models for use in lattice generation. These language models have words implicit within them, though given suitable smoothing can be used to support previously unseen words. In this work, we first present the results of phone and grapheme recognition, in addition to word recognition based on phone and grapheme sub-word units. On the RT04s independent headset microphone (IHM) test condition, we find word error rate (WER) using phone sub-word units lower than that with graphemes, 44.5% compared to 54.5%. The phone error rate (PER) is 48.2%, slightly higher than the grapheme error rate (GER) of 46.3%, though these are not directly comparable as there are fewer graphemes than phones. We then present results on a spoken term detection (STD) task. Again using the RT04s test set, 78 in-vocabulary words and 64 out-of-vocabulary words were selected as search terms from the reference transcription. HTK was used to generate word or sub-word lattices, and a tool developed at Brno [1] used to …

Spoken Language Recognition Based on Gap-Weighted Subsequence Kernels

ViSPer: A Multilingual TTS Approach Based on VITS Using Deep Feature Loss

Speech neuromuscular decoding based on spectrogram images using conformal predictors with Bi-LSTM.

Homogenous Ensemble Phonotactic Language Recognition Based on SVM Supervector Reconstruction

Improved Phonotactic Language Recognition Based on RNN Feature Reconstruction

Language Recognition Based on Acoustic Diversified Phone Recognizers and Phonotactic Feature Fusion

Personalized Speech Recognizer With Keyword-Based Personalized Lexicon And Language Model Using Word Vector Representations

Discriminative Boosting Algorithm for Diversified Front-End Phonotactic Language Recognition

Discriminative Vector Space Model Based Language Recognition

Grapheme-based Spoken Term Detection in the Meetings Domain Extended abstract submitted to MLMI-07

Discriminative Boosting Regression Backend for Phonotactic Language Recognition

Log-Likelihood Kernels Based on Adapted GMMs for Speaker Verification

Speech Enhancement with a GSC-like Structure Employing Sparse Coding

A New Subspace Based Speaker Adaptation Method

Weighted Cluster-Range Loss and Criticality-Enhancement Loss for Speaker Recognition

Phone Lattice Reconstruction for Embedded Language Recognition in LVCSR

Weighted fast sequential DTW for multilingual audio Query-by-Example retrieval

Towards High Performance LVCSR in Speech-to-Speech Translation System on Smart Phones.

Multi-scale Kernels for Short Utterance Speaker Recognition.

Query-by-example Spoken Term Detection Based on Phonetic Posteriorgram

Biologically-Inspired Spike-Based Automatic Speech Recognition of Isolated Digits Over a Reproducing Kernel Hilbert Space