Abstract:External knowledge has been widely applied in image captioning tasks to enrich the generated sentences. However, existing methods retrieve knowledge by considering only semantic relevance while ignoring whether they are useful for captioning. For example, when querying “person” in external knowledge, the most relevant concepts may be “wearing shirt” or “riding horse” statistically, which are not consistent with image contents and introduce noise to generated sentences. Intuitively, we humans can iteratively correlate visual clues with corresponding knowledge to distinguish useful clues from noise. Therefore, we propose an event-aware retrospective learning network for knowledge-based image captioning, which employs a retrospective validation mechanism on captioning models to align the retrieved knowledge with visual contents. This approach is an event-aware perspective and helps select useful knowledge that corresponds to visual facts. To better align images and knowledge, 1) we design an event-aware retrieval algorithm that clusters word-centered knowledge into triplet-centered knowledge (i.e., from “< subject - predicate - object >” to “< triplet A> - edge - < triplet B >”, which provides an event context to facilitate knowledge retrieval and validation. 2) We revisit image contents to retrospectively validate retrieved knowledge by aligning the visual representation between knowledge and image. We summarize the visual characteristics of each knowledge event from the visual genome dataset to help learn which knowledge does not exist in the visual scene and should be discarded. 3) We adopt a dynamic knowledge fusion module that calibrates image and knowledge representations for sentence generation, which includes a knowledge-controlled gate unit that jointly calculates visual and semantic features in event-aware patterns. Compared to current knowledge-based captioning methods, the proposed network retrospectively learns the visual facts by event-aware retrieval and knowledge-image visual alignment, which regularizes the knowledge-incorporated captioning with visual evidence. Extensive experiments on the MS-COCO dataset demonstrate the effectiveness of our method. Ablation studies and visualization demonstrate the advantages of each component of the proposed model.

Attend to Knowledge: Memory-Enhanced Attention Network for Image Captioning.

Image Captioning with Memorized Knowledge

Memory-Augmented Image Captioning

Learning visual relationship and context-aware attention for image captioning

Show, Observe and Tell: Attribute-driven Attention Model for Image Captioning.

Image Caption Generation Method Based on Knowledge Graph Guidance and Self-Attention Mechanism

With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning

An image caption model based on attention mechanism and deep reinforcement learning

Recurrent Image Captioner: Describing Images with Spatial-Invariant Transformation and Attention Filtering

Towards Retrieval-Augmented Architectures for Image Captioning

A Hierarchical Multimodal Attention-based Neural Network for Image Captioning

Learning joint relationship attention network for image captioning

Image captioning with weakly-supervised attention penalty

Bi-Directional Co-Attention Network for Image Captioning

Event-aware Retrospective Learning for Knowledge-based Image Captioning

Combining Object-Based Attention And Attributes For Image Captioning

Image Captioning with a Joint Attention Mechanism by Visual Concept Samples

SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image Captioning

Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive Style