Hybrid Improvements in Multimodal Analysis for Deep Video Understanding

Beibei Zhang,Fan Yu,Yaqun Fang,Tongwei Ren,Gangshan Wu
DOI: https://doi.org/10.1145/3469877.3493599
2021-01-01
Abstract:The Deep Video Understanding Challenge (DVU) is a task that focuses on comprehending long duration videos which involve many entities. Its main goal is to build relationship and interaction knowledge graph between entities to answer relevant questions. In this paper, we improved the joint learning method which we previously proposed in many aspects, including few shot learning, optical flow feature, entity recognition, and video description matching. We verified the effectiveness of these measures through experiments.
What problem does this paper attempt to address?