Audio-Visual Approach For Multimodal Concurrent Speaker Detection

Amit Eliav,Sharon Gannot
2024-07-02
Abstract:Concurrent Speaker Detection (CSD), the task of identifying the presence and overlap of active speakers in an audio signal, is crucial for many audio tasks such as meeting transcription, speaker diarization, and speech separation. This study introduces a multimodal deep learning approach that leverages both audio and visual information. The proposed model employs an early fusion strategy combining audio and visual features through cross-modal attention mechanisms, with a learnable [CLS] token capturing the relevant audio-visual relationships. The model is extensively evaluated on two real-world datasets, AMI and the recently introduced EasyCom dataset. Experiments validate the effectiveness of the multimodal fusion strategy. Ablation studies further support the design choices and the training procedure of the model. As this is the first work reporting CSD results on the challenging EasyCom dataset, the findings demonstrate the potential of the proposed multimodal approach for CSD in real-world scenarios.
Audio and Speech Processing,Image and Video Processing
What problem does this paper attempt to address?
The paper primarily focuses on addressing the problem of Concurrent Speaker Detection (CSD). Specifically, the research aims to identify the presence of active speakers and their overlapping situations in audio signals. This task is crucial for many audio processing applications, such as meeting transcription, speaker diarization, and speech separation. To improve the performance of the CSD task, the authors propose a deep learning-based method that combines audio and visual information. Specifically, the model employs an early fusion strategy, integrating audio and visual features through a cross-modal attention mechanism, and utilizes a learnable [CLS] token to capture relevant audio-visual relationships. The model was extensively evaluated on two real-world datasets—AMI and the recently introduced EasyCom dataset. Experimental results validate the effectiveness of the multimodal fusion strategy, and ablation studies further support the model design choices and training process. Since this is the first study to report CSD results on the challenging EasyCom dataset, the findings suggest that the proposed multimodal approach has potential in real-world CSD tasks. In summary, the main contribution of this paper is the proposal of a multimodal deep learning model for addressing the CSD problem, which effectively combines audio and visual information to improve the accuracy of concurrent speaker detection.