Abstract:External knowledge has been widely applied in image captioning tasks to enrich the generated sentences. However, existing methods retrieve knowledge by considering only semantic relevance while ignoring whether they are useful for captioning. For example, when querying “person” in external knowledge, the most relevant concepts may be “wearing shirt” or “riding horse” statistically, which are not consistent with image contents and introduce noise to generated sentences. Intuitively, we humans can iteratively correlate visual clues with corresponding knowledge to distinguish useful clues from noise. Therefore, we propose an event-aware retrospective learning network for knowledge-based image captioning, which employs a retrospective validation mechanism on captioning models to align the retrieved knowledge with visual contents. This approach is an event-aware perspective and helps select useful knowledge that corresponds to visual facts. To better align images and knowledge, 1) we design an event-aware retrieval algorithm that clusters word-centered knowledge into triplet-centered knowledge (i.e., from “< subject - predicate - object >” to “< triplet A> - edge - < triplet B >”, which provides an event context to facilitate knowledge retrieval and validation. 2) We revisit image contents to retrospectively validate retrieved knowledge by aligning the visual representation between knowledge and image. We summarize the visual characteristics of each knowledge event from the visual genome dataset to help learn which knowledge does not exist in the visual scene and should be discarded. 3) We adopt a dynamic knowledge fusion module that calibrates image and knowledge representations for sentence generation, which includes a knowledge-controlled gate unit that jointly calculates visual and semantic features in event-aware patterns. Compared to current knowledge-based captioning methods, the proposed network retrospectively learns the visual facts by event-aware retrieval and knowledge-image visual alignment, which regularizes the knowledge-incorporated captioning with visual evidence. Extensive experiments on the MS-COCO dataset demonstrate the effectiveness of our method. Ablation studies and visualization demonstrate the advantages of each component of the proposed model.

Image Captioning with Memorized Knowledge

Attend to Knowledge: Memory-Enhanced Attention Network for Image Captioning.

Memory-Augmented Image Captioning

Towards Retrieval-Augmented Architectures for Image Captioning

An image caption model based on attention mechanism and deep reinforcement learning

With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning

Multimodal Memory Modelling for Video Captioning

Image Caption Generation Method Based on Knowledge Graph Guidance and Self-Attention Mechanism

Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive Style

Recurrent Image Captioner: Describing Images with Spatial-Invariant Transformation and Attention Filtering

BENet: bi-directional enhanced network for image captioning

Improving Image Captioning with Better Use of Caption

MeaCap: Memory-Augmented Zero-shot Image Captioning

Intelligent image captioning

Show, Conceive and Tell: Image Captioning with Prospective Linguistic Information

Improving Image Captioning with Better Use of Captions

Neural Image Caption Generation with Weighted Training and Reference

Event-aware Retrospective Learning for Knowledge-based Image Captioning

EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension

Memorial GAN With Joint Semantic Optimization for Unpaired Image Captioning