Abstract:Recent works have shown that prompting large language models with audio encodings can unlock speech recognition capabilities. However, existing techniques do not scale efficiently, especially while handling long form streaming audio inputs -- not only do they extrapolate poorly beyond the audio length seen during training, but they are also computationally inefficient due to the quadratic cost of attention. In this work, we introduce SpeechLLM-XL, a linear scaling decoder-only model for streaming speech recognition. We process audios in configurable chunks using limited attention window for reduced computation, and the text tokens for each audio chunk are generated auto-regressively until an EOS is predicted. During training, the transcript is segmented into chunks, using a CTC forced alignment estimated from encoder output. SpeechLLM-XL with 1.28 seconds chunk size achieves 2.7%/6.7% WER on LibriSpeech test clean/other, and it shows no quality degradation on long form utterances 10x longer than the training utterances.

What problem does this paper attempt to address?

The main issues that this paper attempts to address with existing large language models (SpeechLLMs) when handling long audio inputs are: 1. **Limited length extrapolation capability**: Existing SpeechLLMs show a significant drop in performance when dealing with inputs longer than the maximum audio length encountered during training. This is because these models tend to terminate decoding prematurely when generating transcription text, especially when faced with audio inputs longer than those in the training data. 2. **Low computational efficiency**: The computational cost of the attention mechanism increases quadratically with the length of the audio, leading to high computational overhead when processing long audio inputs. 3. **Latency issues**: Existing non-streaming SpeechLLMs require the entire audio to be received before generating the transcription text, resulting in high perceived latency for users when handling long audio, particularly when deployed in production systems where low latency is a critical requirement. To address these issues, the paper proposes SpeechLLM-XL, a linearly scalable streaming speech recognition model. By segmenting the audio into fixed-length chunks and processing each chunk with a limited attention window, SpeechLLM-XL can effectively handle long audio inputs while maintaining low computational cost and latency. Specifically, the model improves in the following aspects: - **Audio chunking**: The input audio is divided into fixed-length chunks, with each chunk capable of generating a variable number of text tokens. - **Limited attention window**: When processing each audio chunk, the model only focuses on the current chunk and a few preceding chunks, thereby reducing computational overhead. - **Autoregressive generation**: The text tokens for each audio chunk are generated in an autoregressive manner until an end-of-sequence (EOS) token is predicted. Experimental results show that SpeechLLM-XL performs excellently in handling long audio inputs, achieving efficient streaming speech recognition without sacrificing accuracy.

Efficient Streaming LLM for Speech Recognition

Prompting Large Language Models with Speech Recognition Abilities

A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition

Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition

Efficient Streaming Language Models with Attention Sinks

Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time

Faster Speech-LLaMA Inference with Multi-token Prediction

WavLLM: Towards Robust and Adaptive Speech Large Language Model

Decoder-only Architecture for Streaming End-to-end Speech Recognition

Large Language Models Are Strong Audio-Visual Speech Recognition Learners

Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition

Large-scale Language Model Rescoring on Long-form Data

Investigating Decoder-only Large Language Models for Speech-to-text Translation

On decoder-only architecture for speech-to-text and large language model integration

Streaming Long Video Understanding with Large Language Models

Speak While You Think: Streaming Speech Synthesis During Text Generation

SirLLM: Streaming Infinite Retentive LLM

Tuning Large Language Model for Speech Recognition With Mixed-Scale Re-Tokenization

Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding

Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition