Abstract:Video moment localization stands as a crucial task within the realm of computer vision, entailing the identification of temporal moments in untrimmed videos that bear semantic relevance to the supplied natural language queries. This work delves into a relatively unexplored facet of the task: the transferability of video moment localization models. This concern is addressed by evaluating moment localization models within a cross-domain transfer setting. In this setup, we curate multiple datasets distinguished by substantial domain gaps. The model undergoes training on one of these datasets, while validation and testing are executed using the remaining datasets. To confront the challenges inherent in this scenario, we draw inspiration from the recently introduced large-scale pre-trained vision-language models. Our focus is on exploring how the strategic utilization of these resources can bolster the capabilities of a model designed for video moment localization. Nevertheless, the distribution of language queries in video moment localization usually diverges from the text used by pre-trained models, exhibiting distinctions in aspects such as length, content, expression, and more. To mitigate the gap, this work proposes a Moment-Guided Query Prompting (MGQP) method for video moment localization. Our key idea is to generate multiple distinct and complementary prompt primitives through stratification of the original queries. Our approach is comprised of a prompt primitive constructor, a multimodal prompt refiner, and a holistic prompt incorporator. We carry out extensive experiments on Charades-STA, TACoS, DiDeMo, and YouCookII datasets, and investigate the efficacy of the proposed method using various pre-trained models, such as CLIP, ActionCLIP, CLIP4Clip, and VideoCLIP. The experimental results demonstrate the effectiveness of our proposed method.

Transferable Video Moment Localization by Moment-Guided Query Prompting

Cross-modal Moment Localization in Videos.

Prompting Large Language Models to Reformulate Queries for Moment Localization

Attentive Moment Retrieval in Videos

Prompt-based Zero-shot Video Moment Retrieval

The Surprising Effectiveness of Multimodal Large Language Models for Video Moment Retrieval

A Survey on Video Moment Localization

Generative Video Diffusion for Unseen Cross-Domain Video Moment Retrieval

Filling the Information Gap Between Video and Query for Language-Driven Moment Retrieval

Adversarial Video Moment Retrieval by Jointly Modeling Ranking and Localization

MomentDiff: Generative Video Moment Retrieval from Random to Real

Structured Multi-Level Interaction Network for Video Moment Localization via Language Query

Moment Retrieval via Cross-Modal Interaction Networks with Query Reconstruction

Cross-Modal Interaction Networks for Query-Based Moment Retrieval in Videos

DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video

MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval

Context-Enhanced Video Moment Retrieval with Large Language Models

DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval

MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction

Cross-Modal Dynamic Networks for Video Moment Retrieval With Text Query

QD-VMR: Query Debiasing with Contextual Understanding Enhancement for Video Moment Retrieval