Electrolaryngeal Speech Enhancement Based on a Two Stage Framework with Bottleneck Feature Refinement and Voice Conversion

Yaogen Yang,Haozhe Zhang,Zexin Cai,Yao Shi,Ming Li,Dong Zhang,Xiaojun Ding,Jianhua Deng,Jie Wang
DOI: https://doi.org/10.1016/j.bspc.2022.104279
IF: 5.1
2023-01-01
Biomedical Signal Processing and Control
Abstract:An electrolarynx (EL) is a medical device that generates speech for people who lost their biological larynx. However, EL speech signals are unnatural and unintelligible due to the monotonous pitch and the mechanical excitation of the EL device. This paper proposes an end-to-end voice conversion method to enhance EL speech. We adopt a speaker-independent automatic speech recognition model to extract bottleneck features as the intermediate phonetic features for enhancement. Our system includes two stages: the bottleneck feature vectors of the EL speech are mapped by a parallel non-autoregressive model to the corresponding feature vectors of the normal speech in stage one. Then another voice conversion model maps normal speech's bottleneck feature vectors directly to normal speech's Mel-spectrogram in stage two, followed by a MelGAN-based vocoder to convert the Mel-spectrogram into waveform. In addition, we incorporate data augmentation and transfer learning to improve conversion performance. Experimental results show that the proposed method outperforms our baseline methods and performs well in terms of naturalness and intelligibility. The audio samples are available online.(2)
What problem does this paper attempt to address?