On the Role of Visual Cues in Audiovisual Speech Enhancement

Zakaria Aldeneh,Anushree Prasanna Kumar,Barry-John Theobald,Erik Marchi,Sachin Kajarekar,Devang Naik,Ahmed Hussen Abdelaziz
DOI: https://doi.org/10.48550/arXiv.2004.12031
2021-02-25
Abstract:We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show that visual cues provide not only high-level information about speech activity, i.e., speech/silence, but also fine-grained visual information about the place of articulation. One byproduct of this finding is that the learned visual embeddings can be used as features for other visual speech applications. We demonstrate the effectiveness of the learned visual embeddings for classifying visemes (the visual analogy to phonemes). Our results provide insight into important aspects of audiovisual speech enhancement and demonstrate how such models can be used for self-supervision tasks for visual speech applications.
Machine Learning,Computation and Language,Computer Vision and Pattern Recognition,Sound,Audio and Speech Processing
What problem does this paper attempt to address?