First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Tom Tongjia Chen,Hongshan Yu,Zhengeng Yang,Ming Li,Zechuan Li,Jingwen Wang,Wei Miao,Wei Sun,Chen Chen

2023-06-23

Abstract:Affordance-Centric Question-driven Task Completion (AQTC) has been proposed to acquire knowledge from videos to furnish users with comprehensive and systematic instructions. However, existing methods have hitherto neglected the necessity of aligning spatiotemporal visual and linguistic signals, as well as the crucial interactional information between humans and objects. To tackle these limitations, we propose to combine large-scale pre-trained vision-language and video-language models, which serve to contribute stable and reliable multimodal data and facilitate effective spatiotemporal visual-textual alignment. Additionally, a novel hand-object-interaction (HOI) aggregation module is proposed which aids in capturing human-object interaction information, thereby further augmenting the capacity to understand the presented scenario. Our method achieved first place in the CVPR'2023 AQTC Challenge, with a Recall@1 score of 78.7\%. The code is available at <a class="link-external link-https" href="https://github.com/tomchen-ctj/CVPR23-LOVEU-AQTC" rel="external noopener nofollow">this https URL</a>.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The problem this paper attempts to address is: how to acquire knowledge from videos to provide comprehensive and systematic guidance to users, especially in the Affordance-Centric Question-driven Task Completion (AQTC) task. Existing methods have the following shortcomings in handling this task: 1. **Spatiotemporal Alignment Issue**: Existing methods ignore the alignment of visual and language signals in time and space, leading to insufficient understanding and alignment capabilities of multimodal data. 2. **Lack of Human-Object Interaction Information**: Existing methods fail to adequately capture the interaction information between humans and objects, which is particularly important in understanding complex scenes. To address these issues, the authors propose a method that combines large-scale pre-trained vision-language and video-language models, and introduce a new Hand-Object Interaction (HOI) aggregation module to enhance the capture and understanding of interaction information. These improvements enabled the model to achieve first place in the CVPR 2023 AQTC challenge, with a Recall@1 score of 78.7%.

First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

A Solution to CVPR'2023 AQTC Challenge: Video Alignment for Multi-Step Inference

Winning the CVPR'2022 AQTC Challenge: A Two-stage Function-centric Approach

Champion Solution for the WSDM2023 Toloka VQA Challenge

Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models

STVGFormer

Technical Report for CVPR 2022 LOVEU AQTC Challenge

STVGFormer: Spatio-Temporal Video Grounding with Static-Dynamic Cross-Modal Understanding

Unifying 3D Vision-Language Understanding via Promptable Queries

3rd Place Solution for MeViS Track in CVPR 2024 PVUW workshop: Motion Expression guided Video Segmentation

Object Attribute Matters in Visual Question Answering

All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment

Visual Relationship Recognition Via Language And Position Guided Attention

A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter

First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge

1st Place Solution to ECCV 2022 Challenge on Out of Vocabulary Scene Text Understanding: End-to-End Recognition of Out of Vocabulary Words

The Third Place Solution for CVPR2022 AVA Accessibility Vision and Autonomy Challenge

First Place Solution to the ECCV 2024 ROAD++ Challenge @ ROAD++ Spatiotemporal Agent Detection 2024

ObjectNLQ @ Ego4D Episodic Memory Challenge 2024

Cascaded Human-Object Interaction Recognition

Task-driven Visual Saliency and Attention-based Visual Question Answering