Abstract:Videos are generally accompanied with multi-modal information such as audio, text, and motion. The multi-modal information is becoming an important cue for understanding video content. How to model the correlation between multi-modalities in videos is still an unsolved problem in video understanding tasks such as video action recognition, video temporal grounding, and video description. In this talk, we focus on two specific video understanding tasks (i.e., cross-modal self-supervised pretraining and temporal grounding) by exploiting the video-text cross modal information. In particular, we notice that videos are naturally accompanied by abundant text information such as YouTube titles, Instagram captions, and Movie scripts. This textual information could serve as a general information to guide us train a multi-modal network, which could be used as a general video representation to be finetuned on the downstream tasks, or as cross-modal matching similarity to be used for video segment retrieval. Specifically, we first present a general cross-modal pair discrimination (CPD) framework to capture this correlation between a video and its associated text. We train our CPD models on both standard video dataset (Kinetics-210k) and uncurated web video dataset (Instagram-300k) to demonstrate its effectiveness. Without further fine-tuning, the learnt models obtain competitive results for action classification on Kinetics under the linear classification protocol. Moreover, our visual model provides an effective initialization to fine-tune on downstream tasks, which yields a remarkable performance gain for action recognition on UCF101 and HMDB51. Our CPD demonstrates that pre-training on a relatively small dataset is able to yield a comparable performance to those methods of using order magnitude more data, which is meaningful and practicable for the scenarios with limited computational facilities. Second, we present a Contrastive and Compatible Matching Network (C2M-Net), to directly model the relations between language queries and video moments in a joint embedding space. This new metric-learning framework enables fully exploiting negative samples from two new aspects: constructing negative pairs from a dual matching scheme and mining negative pairs across different videos. These new negative samples could enhance the joint representation learning of two modalities via contrastive learning to maximize their mutual information. In addition, to precisely rank relatively positive pairs for accurate temporal grounding, we also learn the compatibility between queries and moments by directly regressing their IoU-based similarity. Our C2M-Net yields state-of-the-art performance on three benchmarks of CharadesSTA, TACoS, and ActivityNet-Captions.

COOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language Representation

Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language Representation Learning and Retrieval

Multimodal Contrastive Training for Visual Representation Learning

Iterative Uni-modal and Cross-modal Clustered Contrastive Learning for Image-text Retrieval

COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval

Cross-modal Pretraining and Matching for Video Understanding

Cross-modality interaction reasoning for enhancing vision-language pre-training in image-text retrieval

Vision Learners Meet Web Image-Text Pairs

Transformer Decoders with MultiModal Regularization for Cross-Modal Food Retrieval

Vision-Language Pre-Training with Triple Contrastive Learning

Knowledge Boosting: Rethinking Medical Contrastive Vision-Language Pre-Training

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval

Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning

ProtoCLIP: Prototypical Contrastive Language Image Pretraining

GrowCLIP: Data-aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-training

ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

Efficient Token-Guided Image-Text Retrieval With Consistent Multimodal Contrastive Training

Improving Cross-Modal Image-Text Retrieval With Teacher-Student Learning

Iclip: Bridging Image Classification and Contrastive Language-Image Pre-Training for Visual Recognition