Abstract:Large-scale pre-trained multi-modal models (e.g., CLIP) demonstrate strong zero-shot transfer capability in many discriminative tasks. Their adaptation to zero-shot image-conditioned text generation tasks has drawn increasing interest. Prior arts approach to zero-shot captioning by either utilizing the existing large language models (e.g., GPT-2) or pre-training the encoder-decoder network in an end-to-end manner. In this work, we propose a simple framework, named DeCap, for zero-shot captioning. We introduce a lightweight visual-aware language decoder. This decoder is both data-efficient and computation-efficient: 1) it only requires the text data for training, easing the burden on the collection of paired data. 2) it does not require end-to-end training. When trained with text-only data, the decoder takes the text embedding extracted from the off-the-shelf CLIP encoder as a prefix embedding. The challenge is that the decoder is trained on the text corpus but at the inference stage, it needs to generate captions based on visual inputs. The modality gap issue is widely observed in multi-modal contrastive models that prevents us from directly taking the visual embedding as the prefix embedding. We propose a training-free mechanism to reduce the modality gap. We project the visual embedding into the CLIP text embedding space, while the projected embedding retains the information of the visual input. Taking the projected embedding as the prefix embedding, the decoder generates high-quality descriptions that match the visual input. The experiments show that DeCap outperforms other zero-shot captioning methods and unpaired captioning methods on the typical image captioning benchmarks, i.e., MSCOCO and NoCaps.

CLIP4Caption ++: Multi-CLIP for Video Caption

CLIP4Caption: CLIP for Video Caption

SBAT: Video Captioning with Sparse Boundary-Aware Transformer

Cap4Video++: Enhancing Video Understanding with Auxiliary Captions

Vatex Video Captioning Challenge 2020: Multi-View Features and Hybrid Reward Strategies for Video Captioning

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

A CLIP-Enhanced Method for Video-Language Understanding

MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning

Multi-scale features with temporal information guidance for video captioning

Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?

ClipCap: CLIP Prefix for Image Captioning

Based-CLIP early fusion transformer for image caption

End-to-End 3D Dense Captioning with Vote2Cap-DETR

FiGCLIP: Fine-Grained CLIP Adaptation via Densely Annotated Videos

CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning

DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Delving Deeper into the Decoder for Video Captioning

From Scarcity to Efficiency: Improving CLIP Training via Visual-enriched Captions

Fine-grained Image Captioning with CLIP Reward

A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions