Hallucination in Perceptual Metric-Driven Speech Enhancement Networks

George Close,Thomas Hain,Stefan Goetze
2024-05-24
Abstract:Within the area of speech enhancement, there is an ongoing interest in the creation of neural systems which explicitly aim to improve the perceptual quality of the processed audio. In concert with this is the topic of non-intrusive (i.e. without clean reference) speech quality prediction, for which neural networks are trained to predict human-assigned quality labels directly from distorted audio. When combined, these areas allow for the creation of powerful new speech enhancement systems which can leverage large real-world datasets of distorted audio, by taking inference of a pre-trained speech quality predictor as the sole loss function of the speech enhancement system. This paper aims to identify a potential pitfall with this approach, namely hallucinations which are introduced by the enhancement system `tricking' the speech quality predictor.
Sound,Audio and Speech Processing
What problem does this paper attempt to address?
The paper attempts to address the issue of "hallucination" phenomena in perceptual metric-driven speech enhancement networks. Specifically, when using non-intrusive (i.e., no need for clean reference signals) speech quality predictors to optimize speech enhancement systems, these systems may "trick" the quality predictors by introducing some distortions or hallucinations that are not real but can improve the predictor's score. Although these hallucinations can lead to higher scores from the predictors, they significantly degrade the quality of speech in actual human auditory evaluations. The main objectives of the paper include: 1. **Understanding how speech enhancement systems learn to optimize their output to fit non-intrusive quality predictors during training**: Investigating how these systems improve the predictor's score through specific distortions or hallucinations. 2. **Identifying why non-intrusive quality predictors fail to accurately assess the actual human perceived quality of speech enhancement systems**: Exploring why these predictors sometimes produce results inconsistent with human evaluations, even when the speech enhancement systems are not directly optimized for these predictors. Through these studies, the paper aims to reveal potential issues when using non-intrusive quality predictors for speech enhancement and provide guidance for future research and applications.