DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Wei Li,Linchao Zhu,Longyin Wen,Yi Yang

2023-03-06

Abstract:Large-scale pre-trained multi-modal models (e.g., CLIP) demonstrate strong zero-shot transfer capability in many discriminative tasks. Their adaptation to zero-shot image-conditioned text generation tasks has drawn increasing interest. Prior arts approach to zero-shot captioning by either utilizing the existing large language models (e.g., GPT-2) or pre-training the encoder-decoder network in an end-to-end manner. In this work, we propose a simple framework, named DeCap, for zero-shot captioning. We introduce a lightweight visual-aware language decoder. This decoder is both data-efficient and computation-efficient: 1) it only requires the text data for training, easing the burden on the collection of paired data. 2) it does not require end-to-end training. When trained with text-only data, the decoder takes the text embedding extracted from the off-the-shelf CLIP encoder as a prefix embedding. The challenge is that the decoder is trained on the text corpus but at the inference stage, it needs to generate captions based on visual inputs. The modality gap issue is widely observed in multi-modal contrastive models that prevents us from directly taking the visual embedding as the prefix embedding. We propose a training-free mechanism to reduce the modality gap. We project the visual embedding into the CLIP text embedding space, while the projected embedding retains the information of the visual input. Taking the projected embedding as the prefix embedding, the decoder generates high-quality descriptions that match the visual input. The experiments show that DeCap outperforms other zero-shot captioning methods and unpaired captioning methods on the typical image captioning benchmarks, i.e., MSCOCO and NoCaps.

Computer Vision and Pattern Recognition,Artificial Intelligence,Computation and Language

What problem does this paper attempt to address?

The problem that this paper attempts to solve is to generate automatic descriptions from images to text without the need for human - annotated data, that is, the zero - shot image captioning task. Specifically, the author proposes a new framework - DeCap, which aims to generate high - quality image captions by training a lightweight visual - perception - language decoder using only text data. This method solves the dependence on paired image - text data in existing methods and the problem of huge computational resources required for end - to - end pre - training on large - scale network data. In addition, DeCap also proposes a training - free mechanism to reduce the gap between modalities, thereby making it possible to generate captions from visual inputs. This method not only improves the quality of zero - shot image captioning, but also enhances the generalization ability of the model in different captioning tasks.

DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

DRCap: Decoding CLAP Latents with Retrieval-augmented Generation for Zero-shot Audio Captioning

MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning

MeaCap: Memory-Augmented Zero-shot Image Captioning

Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training

ClipCap: CLIP Prefix for Image Captioning

Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment

Transferable Decoding with Visual Entities for Zero-Shot Image Captioning

Zero-TextCap: Zero-shot Framework for Text-based Image Captioning

Delving Deeper into the Decoder for Video Captioning

Entrocap: Zero-Shot Image Captioning with Entropy-Based Retrieval

CLIP4Caption: CLIP for Video Caption

Zero-Shot Audio Captioning Using Soft and Hard Prompts

From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping

CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification

HyperCLIP: Adapting Vision-Language models with Hypernetworks

Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss

Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm