Speech Synthesis from Electrocorticogram During Imagined Speech Using a Transformer-Based Decoder and Pretrained Vocoder

Shuji Komeiji,Kai Shigemi,Takumi Mitsuhashi,Yasushi Iimura,Hiroharu Suzuki,Hidenori Sugano,Koich Shinoda,Kohei Yatabe,Toshihisa Tanaka
DOI: https://doi.org/10.1101/2024.08.21.608927
2024-08-22
Abstract:This study describes speech synthesis from an Electrocorticogram (ECoG) during imagined speech. We aim to generate high-quality audio despite the limitations of available training data by employing a Transformer-based decoder and a pretrained vocoder. Specifically, we used a pre-trained neural vocoder, Parallel WaveGAN, to convert the log-mel spectrograms output by the Transformer decoder, which was trained on ECoG signals, into high-quality audio signals. In our experiments, using ECoG signals recorded from 13 participants, the synthesized speech from imagined speech achieved a dynamic time-warping (DTW) Pearson correlation ranging from 0.85 to 0.95. This high-quality speech synthesis can be attributed to the Transformer decoder's ability to accurately reconstruct high-fidelity log-mel spectrograms, demonstrating its effectiveness in dealing with limited training data.
Neuroscience
What problem does this paper attempt to address?