Abstract:End-to-end (E2E) automatic speech recognition (ASR) models have become standard practice for various commercial applications. However, in real-world scenarios, the long-tailed nature of word distribution often leads E2E ASR models to perform well on common words but fall short in recognizing uncommon ones. Recently, the notion of a contextual adapter (CA) was proposed to infuse external knowledge represented by a context word list into E2E ASR models. Although CA can improve recognition performance on rare words, two crucial data imbalance problems remain. First, when using low-frequency words as context words during training, since these words rarely occur in the utterance, CA becomes prone to overfit on attending to the <no-context> token due to higher-frequency words not being present in the context list. Second, the long-tailed distribution within the context list itself still causes the model to perform poorly on low-frequency context words. In light of this, we explore in-depth the impact of altering the context list to have words with different frequency distributions on model performance, and meanwhile extend CA with a simple yet effective context-balanced learning objective. A series of experiments conducted on the AISHELL-1 benchmark dataset suggests that using all vocabulary words from the training corpus as the context list and pairing them with our balanced objective yields the best performance, demonstrating a significant reduction in character error rate (CER) by up to 1.21% and a more pronounced 9.44% reduction in the error rate of zero-shot words.

CASA-ASR: Context-Aware Speaker-Attributed ASR.

CASA Based Speech Separation for

A Comparative Study on Multichannel Speaker-Attributed Automatic Speech Recognition in Multi-party Meetings

CASA Based Speech Separation for Robust Speech Recognition

Integrating Source-Channel and Attention-Based Sequence-to-Sequence Models for Speech Recognition

End-to-end Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis

SA-Paraformer: Non-autoregressive End-to-End Speaker-Attributed ASR

Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation

A Comparative Study on Speaker-attributed Automatic Speech Recognition in Multi-party Meetings.

Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications

Speaker conditioned acoustic modeling for multi-speaker conversational ASR

An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition

Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment

Robust Front-End for Speech Recognition Based on Computational Auditory Scene Analysis and Speaker Model

Space-and-speaker-aware Acoustic Modeling with Effective Data Augmentation for Recognition of Multi-Array Conversational Speech

SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization Ability

MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction

Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring

SAN: a Robust End-to-end ASR Model Architecture

Deep CLAS: Deep Contextual Listen, Attend and Spell