Deep Video Understanding with Video-Language Model

Runze Liu,Yaqun Fang,Fan Yu,Ruiqi Tian,Tongwei Ren,Gangshan Wu
DOI: https://doi.org/10.1145/3581783.3612863
2023-01-01
Abstract:Pre-trained video-language models (VLMs) have shown superior performance in high-level video understanding tasks, analyzing multi-modal information, aligning with Deep Video Understanding Challenge (DVUC) requirements. In this paper, we explore pretrained VLMs' potential in multimodal question answering for longform videos. We propose a solution called Dual Branches Video Modeling (DBVM), which combines knowledge graph (KG) and VLMs, leveraging their strengths and addressing shortcomings. The KG branch recognizes and localizes entities, fuses multimodal features at different levels, and constructs KGs with entities as nodes and relationships as edges. The VLM branch applies a selection strategy to adapt input movies into acceptable length and a cross-matching strategy to post-process results providing accurate scene descriptions. Experiments conducted on the DVUC dataset validate the effectiveness of our DBVM.
What problem does this paper attempt to address?