Abstract:External knowledge has been widely applied in image captioning tasks to enrich the generated sentences. However, existing methods retrieve knowledge by considering only semantic relevance while ignoring whether they are useful for captioning. For example, when querying “person” in external knowledge, the most relevant concepts may be “wearing shirt” or “riding horse” statistically, which are not consistent with image contents and introduce noise to generated sentences. Intuitively, we humans can iteratively correlate visual clues with corresponding knowledge to distinguish useful clues from noise. Therefore, we propose an event-aware retrospective learning network for knowledge-based image captioning, which employs a retrospective validation mechanism on captioning models to align the retrieved knowledge with visual contents. This approach is an event-aware perspective and helps select useful knowledge that corresponds to visual facts. To better align images and knowledge, 1) we design an event-aware retrieval algorithm that clusters word-centered knowledge into triplet-centered knowledge (i.e., from “< subject - predicate - object >” to “< triplet A> - edge - < triplet B >”, which provides an event context to facilitate knowledge retrieval and validation. 2) We revisit image contents to retrospectively validate retrieved knowledge by aligning the visual representation between knowledge and image. We summarize the visual characteristics of each knowledge event from the visual genome dataset to help learn which knowledge does not exist in the visual scene and should be discarded. 3) We adopt a dynamic knowledge fusion module that calibrates image and knowledge representations for sentence generation, which includes a knowledge-controlled gate unit that jointly calculates visual and semantic features in event-aware patterns. Compared to current knowledge-based captioning methods, the proposed network retrospectively learns the visual facts by event-aware retrieval and knowledge-image visual alignment, which regularizes the knowledge-incorporated captioning with visual evidence. Extensive experiments on the MS-COCO dataset demonstrate the effectiveness of our method. Ablation studies and visualization demonstrate the advantages of each component of the proposed model.

Event-aware Retrospective Learning for Knowledge-based Image Captioning

SKANet - Structured Knowledge-Aware Network for Visual Dialog.

Attend to Knowledge: Memory-Enhanced Attention Network for Image Captioning.

Image Caption Generation Method Based on Knowledge Graph Guidance and Self-Attention Mechanism

Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training Model

Caption Matters: a New Perspective for Knowledge-Based Visual Question Answering

Image Captioning with Memorized Knowledge

Towards Retrieval-Augmented Architectures for Image Captioning

Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning

Enhancing Video Event Recognition Using Automatically Constructed Semantic-Visual Knowledge Base.

K-LITE: Learning Transferable Visual Models with External Knowledge

Context-Aware Visual Policy Network for Fine-Grained Image Captioning

You should know more: Learning external knowledge for visual dialog

General Knowledge Embedded Image Representation Learning

Memory-Augmented Image Captioning

Learning visual relationship and context-aware attention for image captioning

Region-Aware Image Captioning Via Interaction Learning

Show, Conceive and Tell: Image Captioning with Prospective Linguistic Information

Motion-Aware Video Paragraph Captioning Via Exploring Object-Centered Internal Knowledge

Adaptively Attending to Visual Attributes and Linguistic Knowledge for Captioning