Abstract:External knowledge has been widely applied in image captioning tasks to enrich the generated sentences. However, existing methods retrieve knowledge by considering only semantic relevance while ignoring whether they are useful for captioning. For example, when querying “person” in external knowledge, the most relevant concepts may be “wearing shirt” or “riding horse” statistically, which are not consistent with image contents and introduce noise to generated sentences. Intuitively, we humans can iteratively correlate visual clues with corresponding knowledge to distinguish useful clues from noise. Therefore, we propose an event-aware retrospective learning network for knowledge-based image captioning, which employs a retrospective validation mechanism on captioning models to align the retrieved knowledge with visual contents. This approach is an event-aware perspective and helps select useful knowledge that corresponds to visual facts. To better align images and knowledge, 1) we design an event-aware retrieval algorithm that clusters word-centered knowledge into triplet-centered knowledge (i.e., from “< subject - predicate - object >” to “< triplet A> - edge - < triplet B >”, which provides an event context to facilitate knowledge retrieval and validation. 2) We revisit image contents to retrospectively validate retrieved knowledge by aligning the visual representation between knowledge and image. We summarize the visual characteristics of each knowledge event from the visual genome dataset to help learn which knowledge does not exist in the visual scene and should be discarded. 3) We adopt a dynamic knowledge fusion module that calibrates image and knowledge representations for sentence generation, which includes a knowledge-controlled gate unit that jointly calculates visual and semantic features in event-aware patterns. Compared to current knowledge-based captioning methods, the proposed network retrospectively learns the visual facts by event-aware retrieval and knowledge-image visual alignment, which regularizes the knowledge-incorporated captioning with visual evidence. Extensive experiments on the MS-COCO dataset demonstrate the effectiveness of our method. Ablation studies and visualization demonstrate the advantages of each component of the proposed model.

Towards Retrieval-Augmented Architectures for Image Captioning

Towards Retrieval-Augmented Architectures for Image Captioning

Image Captioning with Memorized Knowledge

Memory-Augmented Image Captioning

Attend to Knowledge: Memory-Enhanced Attention Network for Image Captioning.

With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning

EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension

Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning

Recurrent Image Captioner: Describing Images with Spatial-Invariant Transformation and Attention Filtering

No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning

An image caption model based on attention mechanism and deep reinforcement learning

Retrieval-Augmented Multimodal Language Modeling

An Ensemble of Generation- and Retrieval-Based Image Captioning With Dual Generator Generative Adversarial Network

BENet: bi-directional enhanced network for image captioning

Event-aware Retrospective Learning for Knowledge-based Image Captioning

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding

Improving Image Captioning with Better Use of Caption

Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage

A Study of ConvNeXt Architectures for Enhanced Image Captioning

Incorporating retrieval-based method for feature enhanced image captioning

Exploring Visual Relationship for Image Captioning