A Joint Detection-Classification Model for Weakly Supervised Sound Event Detection Using Multi-Scale Attention Method

Yaoguang Wang,Liang He
DOI: https://doi.org/10.1109/isspit51521.2020.9408948
2020-01-01
Abstract:Attention mechanism has been applied to the weakly supervised sound event detection (SED) and has achieved state-of-the-art performance, but most methods only concentrate along the time axis. In this paper, we propose the multi-scale time-frequency attention (MTFA) method to capture the intrinsic features at different scales both in time and frequency domain for audio tagging (AT) and SED. Our model is a unified network which can perform AT and SED simultaneously, it produces multi-scale attention-aware representations for SED with MTFA module, and a global pooling module maps the representations to presence probability of corresponding audio event for AT. To evaluate the proposed method, we conduct experiments on Task4 of Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, and it achieves 57.9% (F1-score) in AT task and 0.71 (error rate) in SED task on evaluation set, which is comparable to the state-of-the-art results in the challenge.
What problem does this paper attempt to address?