Abstract:Objective. Smart hearing aids which can decode the focus of a user's attention could considerably improve comprehension levels in noisy environments. Methods for decoding auditory attention from electroencephalography (EEG) have attracted considerable interest for this reason. Recent studies suggest that the integration of deep neural networks (DNNs) into existing auditory attention decoding algorithms is highly beneficial, although it remains unclear whether these enhanced algorithms can perform robustly in different real-world scenarios. To this end, we sought to characterise the performance of DNNs at reconstructing the envelope of an attended speech stream from EEG recordings in different listening conditions. In addition, given the relatively sparse availability of EEG data, we investigate possibility of applying subject-independent algorithms to EEG recorded from unseen individuals. Approach. Both linear models and nonlinear DNNs were employed to decode the envelope of clean speech from EEG recordings, with and without subject-specific information. The mean behaviour, as well as the variability of the reconstruction, was characterised for each model. We then trained subject-specific linear models and DNNs to reconstruct the envelope of speech in clean and noisy conditions, and investigated how well they performed in different listening scenarios. We also established that these models can be used to decode auditory attention in competing-speaker scenarios. Main results. The DNNs offered a considerable advantage over their linear counterpart at reconstructing the envelope of clean speech. This advantage persisted even when subject-specific information was unavailable at the time of training. The same DNN architectures generalised to a distinct dataset, which contained EEG recorded under a variety of listening conditions. In competing-speakers and speech-in-noise conditions, the DNNs significantly outperformed the linear models. Finally, the DNNs offered a considerable improvement over the linear approach at decoding auditory attention in competing-speakers scenarios. Significance. We present the first detailed study into the extent to which DNNs can be employed for reconstructing the envelope of an attended speech stream. We conclusively demonstrate that DNNs have the ability to improve the reconstruction of the attended speech envelope. The variance of the reconstruction error is shown to be similar for both DNNs and the linear model. Overall, DNNs are demonstrated to show promise for real-world auditory attention decoding, since they perform well in multiple listening conditions and generalise to data recorded from unseen participants.

Predicting EEG Responses to Attended Speech via Deep Neural Networks for Speech

Identification of Attended Speech Stream Using Single-Trial Electroencephalography Recording

Robust decoding of the speech envelope from EEG recordings through deep neural networks

Deep learning-based auditory attention decoding in listeners with hearing impairment

EEG-Based Short-Time Auditory Attention Detection Using Multi-Task Deep Learning.

Linear versus deep learning methods for noisy speech separation for EEG-informed attention decoding

Comparison of linear and nonlinear methods for decoding selective attention to speech from ear-EEG recordings

Dissecting neural computations in the human auditory pathway using deep neural networks for speech

Deep Neural Network Based Noised Asian Speech Enhancement and Its Implementation on a Hearing Aid App.

Recognition of words from brain-generated signals of speech-impaired people: Application of autoencoders as a neural Turing machine controller in deep neural networks

A model of speech recognition for hearing-impaired listeners based on deep learning

Hearing-Loss Compensation Using Deep Neural Networks: A Framework and Results From a Listening Test

Extracting the Auditory Attention in a Dual-Speaker Scenario From EEG Using a Joint CNN-LSTM Model

A Novel Deep Learning Architecture for Decoding Imagined Speech from EEG

Dynamic modeling of EEG responses to natural speech reveals earlier processing of predictable words

Prediction of speech intelligibility with DNN-based performance measures

Identification of perceived sentences using deep neural networks in EEG

NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

EEG-informed attended speaker extraction from recorded speech mixtures with application in neuro-steered hearing prostheses

Exploiting Hidden Representations from a DNN-based Speech Recogniser for Speech Intelligibility Prediction in Hearing-impaired Listeners

Efficient Automatic Speech Recognition from EEG Signals Using Optimal Deep Learning Approach