Abstract:Although keyword spotting (KWS) technologies have been successfully applied to some applications, most KWS systems have a common problem of noise-robustness when applied to real-world environments. Audio-visual keyword spotting (AVKWS) using both acoustic and visual information is a solution to complementarily solve the problem. Most existing audio-visual speech recognition (AVSR) systems extract geometric features as visual features, which heavily rely on accurate and reliable detection and tracking of facial feature points. To avoid this defect of geometric features, an appearance-based discriminative local spatial-temporal descriptor (disCLBP-TOP) is proposed in this paper, which devotes to extracting robust and discriminative patterns of interest. Besides, a parallel two-step recognition based on both acoustic and visual keyword searching and re-scoring is conducted, which complementarily makes the best of two modalities under different noisy conditions. Adaptive weights for decision fusion are generated using a sigmoid function based on reliabilities of the two modalities, capable of adapting to various noisy conditions. Experiments show that our proposed parallel AVKWS strategy based on decision fusion significantly improves the noise robustness and attains better performance than feature fusion based audio-visual spotter. Additionally, disCLBP-TOP shows more competitive performance than CLBP-TOP.

Keyword-specific normalization based keyword spotting for spontaneous speech

A New Keyword Spotting Approach for Spontaneous Mandarin Speech

Keyword Spotting Based on Phoneme Confusion Matrix

A VOCABULARY-INDEPENDENT KEYWORD SPOTTER FOR SPONTANEOUS CHINESE SPEECH

Keyword-Specific Acoustic Model Pruning for Open-Vocabulary Keyword Spotting

A Keyword Spotting Method

PhonMatchNet: Phoneme-Guided Zero-Shot Keyword Spotting for User-Defined Keywords

Audio-visual Keyword Spotting for Mandarin Based on Discriminative Local Spatial-Temporal Descriptors.

Open-vocabulary Keyword-spotting with Adaptive Instance Normalization

Keyword Spotting Based on Hypothesis Boundary Realignment and State-Level Confidence Weighting

MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting

Keyword Spotting Based on Syllable Confusion Network.

An Approach of Keyword Spotting Based on HMM

Encoder-Decoder Neural Architecture Optimization for Keyword Spotting

HarkMan—A Vocabulary-Independent Keyword Spotter for Spontaneous Chinese Speech

TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer

Audio-visual Keyword Spotting Based on Adaptive Decision Fusion under Noisy Conditions for Human-Robot Interaction.

LEXICAL ACCESS-BASED CONFIDENCE MEASURE FOR A SPANISH KEYWORD SPOTTING SYSTEM

U2-KWS: Unified Two-pass Open-vocabulary Keyword Spotting with Keyword Bias

Keyword spotting -- Detecting commands in speech using deep learning

Convolutional Recurrent Neural Networks for Small-Footprint Keyword Spotting