Abstract:Recent progress in video-text retrieval has been driven largely by advancements in model architectures and training strategies. However, the representation learning capabilities of videotext retrieval models remain constrained by lowquality and limited training data annotations. To address this issue, we present a novel ViDeoText Retrieval Paradigm with RElevance-based AugMentation, namely DREAM, which enhances video and text data using large foundation models to learn more generalized features. Specifically, we first adopt a simple augmentation method, which generates self-similar data by randomly duplicating or dropping subwords and frames. In addition, inspired by the recent advancement in visual and language generative models, we propose a more robust augmentation method through textual paraphrasing and video stylization using large language models (LLMs) and visual generative models (VGMs). To further enrich video and text information, we propose a relevance-based augmentation method, where LLMs and VGMs generate and integrate new relevant information into the original data. Leveraging this enriched data, extensive experiments on several video-text retrieval benchmarks demonstrate the superiority of DREAM over existing methods.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is the limited representation learning ability in video - text retrieval (VTR) due to low - quality and limited training data annotation. Specifically, existing VTR models rely on one - to - one linearly annotated datasets during training, and these datasets may have inaccurate or ambiguous annotations, which thus affect the learning effect and generalization ability of the models. To solve this problem, the author proposes a new method named DREAM, which improves the quality and diversity of video and text data through relevance - based augmentation. DREAM utilizes large - scale foundation models (such as large language models LLMs and visual generation models VGMs) to generate more diverse video and text data to help the model learn more general features. ### Main Contributions 1. **Identifying Challenges**: The paper points out that the main challenge in video - text retrieval lies in the ambiguous one - to - one labels, which hinder the model from learning robust representations. Therefore, the author explores data augmentation techniques in multiple dimensions to address this challenge. 2. **Proposing the DREAM Framework**: DREAM includes three augmentation methods: - Simple Augmentation (SA): Generate semantically similar videos and texts by randomly copying or deleting frames and words. - Augmentation by Text Paraphrasing and Video Stylization (TPVS): Use off - the - shelf large - scale language models and visual generation models to generate semantically similar videos and texts. - Augmentation by Relevance Enhancing (RE): Enrich video and text content by introducing additional information related to the original data. 3. **Experimental Verification**: Extensive experiments show that DREAM significantly outperforms existing methods on three popular benchmark datasets, MSR - VTT, MSVD, and ActivityNet, achieving state - of - the - art performance. ### Experimental Results - **MSR - VTT**: DREAM significantly surpasses all previous methods in multiple retrieval metrics. For example, Recall@1 reaches 60.8 and 60.6, Recall@5 reaches 84.5 and 85.2, and Recall@5 reaches 91.4 and 92.5. - **MSVD**: DREAM achieves Recall@1 of 61.6, Recall@5 of 87.1, and Recall@10 of 93.2 in Text - to - Video retrieval. - **ActivityNet**: DREAM achieves the highest score of Recall@1 of 59.1 in Text - to - Video retrieval. Through these improvements, DREAM not only improves the performance of video - text retrieval but also provides new directions for future research.

DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation

VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model

UATVR: Uncertainty-Adaptive Text-Video Retrieval

Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Reading-strategy Inspired Visual Representation Learning for Text-to-Video Retrieval

Distilling Vision-Language Models on Millions of Videos

Mimir: Improving Video Diffusion Models for Precise Text Understanding

Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension

Utilizing Text-based Augmentation to Enhance Video Captioning

Generative Video Diffusion for Unseen Cross-Domain Video Moment Retrieval

Enhancing Long Video Understanding via Hierarchical Event-Based Memory

Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-Training

A Multi-interaction Model with Cross-Branch Feature Fusion for Video-Text Retrieval.

Memory Enhanced Embedding Learning for Cross-Modal Video-Text Retrieval

Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation

LLM4VG: Large Language Models Evaluation for Video Grounding

VEnhancer: Generative Space-Time Enhancement for Video Generation

VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning