Abstract:Video moment localization stands as a crucial task within the realm of computer vision, entailing the identification of temporal moments in untrimmed videos that bear semantic relevance to the supplied natural language queries. This work delves into a relatively unexplored facet of the task: the transferability of video moment localization models. This concern is addressed by evaluating moment localization models within a cross-domain transfer setting. In this setup, we curate multiple datasets distinguished by substantial domain gaps. The model undergoes training on one of these datasets, while validation and testing are executed using the remaining datasets. To confront the challenges inherent in this scenario, we draw inspiration from the recently introduced large-scale pre-trained vision-language models. Our focus is on exploring how the strategic utilization of these resources can bolster the capabilities of a model designed for video moment localization. Nevertheless, the distribution of language queries in video moment localization usually diverges from the text used by pre-trained models, exhibiting distinctions in aspects such as length, content, expression, and more. To mitigate the gap, this work proposes a Moment-Guided Query Prompting (MGQP) method for video moment localization. Our key idea is to generate multiple distinct and complementary prompt primitives through stratification of the original queries. Our approach is comprised of a prompt primitive constructor, a multimodal prompt refiner, and a holistic prompt incorporator. We carry out extensive experiments on Charades-STA, TACoS, DiDeMo, and YouCookII datasets, and investigate the efficacy of the proposed method using various pre-trained models, such as CLIP, ActionCLIP, CLIP4Clip, and VideoCLIP. The experimental results demonstrate the effectiveness of our proposed method.

Language Guided Networks for Cross-modal Moment Retrieval

Cross-modal Moment Localization in Videos.

Attentive Moment Retrieval in Videos

Cross-Modal Video Moment Retrieval with Spatial and Language-Temporal Attention.

Dual-Channel Localization Networks for Moment Retrieval with Natural Language

Cross-Modal Interaction Network for Video Moment Retrieval.

Moment Retrieval via Cross-Modal Interaction Networks with Query Reconstruction

Cross-Modal Dynamic Networks for Video Moment Retrieval With Text Query

Cross-Modality Knowledge Calibration Network for Video Corpus Moment Retrieval

Moment is Important: Language-Based Video Moment Retrieval Via Adversarial Learning

Transferable Video Moment Localization by Moment-Guided Query Prompting

Semantic Collaborative Learning for Cross-Modal Moment Localization

A Multi-modal Fusion Algorithm for Cross-modal Video Moment Retrieval

Cross-Modal Interaction Networks for Query-Based Moment Retrieval in Videos

Coarse-to-Fine Semantic Alignment for Cross-Modal Moment Localization

Structured Multi-Level Interaction Network for Video Moment Localization via Language Query

Multi-stage Aggregated Transformer Network for Temporal Language Localization in Videos

Collaborative Spatial-Temporal Interaction for Language-Based Moment Retrieval

Weakly Supervised Moment Localization with Natural Language Based on Semantic Reconstruction

You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in Videos

MLN: Moment localization Network and Samples Selection for Moment Retrieval