Spatial attentional bilinear 3d convolutional network for video-based autism spectrum disorder detection

Kangbo Sun,Lin Li,Lianqiang Li,Ningyu He,Jie Zhu
DOI: https://doi.org/10.1109/icassp40776.2020.9054641
2020-01-01
Abstract:Video-based Autism Spectrum Disorder (ASD) detection is a challenge to most video classification networks due to the high degree of similarity between categories. Bilinear pooling is a second-order method, which is widely used in finegrained visual recognition. However, the average summation in bilinear pooling limits its ability to perceive spatial information, which is detrimental to fine-grained visual recognition. In this paper, we propose spatial attentional bilinear pooling to enhance its spatial information extraction without significantly increasing the parameters. Further, we propose a fine-grained action recognition network named SA-B3D with LSTM model for video-based ASD detection. The proposed model can focus on more discriminative regions dynamically and effectively. Compared with state-of-the-art models, the proposed model achieves significant improvement on video-based ASD dataset.
What problem does this paper attempt to address?