Exploiting Auxiliary Caption for Video Grounding

Hongxiang Li,Meng Cao,Xuxin Cheng,Yaowei Li,Zhihong Zhu,Yuexian Zou
DOI: https://doi.org/10.1609/aaai.v38i17.29812
2024-01-01
Abstract:Video grounding aims to locate a moment of interest matching the given querysentence from an untrimmed video. Previous works ignore the sparsity dilemmain video annotations, which fails to provide the context information betweenpotential events and query sentences in the dataset. In this paper, we contendthat exploiting easily available captions which describe general actions, i.e.,auxiliary captions defined in our paper, will significantly boost theperformance. To this end, we propose an Auxiliary Caption Network (ACNet) forvideo grounding. Specifically, we first introduce dense video captioning togenerate dense captions and then obtain auxiliary captions by Non-AuxiliaryCaption Suppression (NACS). To capture the potential information in auxiliarycaptions, we propose Caption Guided Attention (CGA) project the semanticrelations between auxiliary captions and query sentences into temporal spaceand fuse them into visual representations. Considering the gap betweenauxiliary captions and ground truth, we propose Asymmetric Cross-modalContrastive Learning (ACCL) for constructing more negative pairs to maximizecross-modal mutual information. Extensive experiments on three public datasets(i.e., ActivityNet Captions, TACoS and ActivityNet-CG) demonstrate that ourmethod significantly outperforms state-of-the-art methods.
What problem does this paper attempt to address?