Abstract:Image captioning aims at understanding various semantic concepts (e.g., objects and relationships) from an image and integrating them in a sentence-level description. Hence, it is necessary to learn the interaction among these concepts. If we define the context of the interaction to be involved in the subject-predicate-object triplet, most current methods only focus on the single triplet for the first-order interaction to generate sentences. Intuitively, we humans are able to perceive the high-order interaction among concepts from two or more triplets to describe an image. For example, when we see the triplets man-cutting-sandwich and man-with-knife, it is natural to integrate and predict the sentence man cutting sandwich with knife. This depends on the high-order interaction between cutting and knife in different triplets. Therefore, exploiting high-order interaction is expected to benefit image captioning and focus on reasoning. In this paper, we introduce the novel high-order interaction learning method over detected objects and relationships for image captioning under the umbrella of the encoder-decoder framework. We first extract a set of object and relationship features in an image. During the encoding stage, the interactive refining network is proposed to learn high-order representations by modeling intra- and inter-object feature interaction in the self-attention fashion. During the decoding stage, the interactive fusion network is proposed to integrate object and relationship information by strengthening their high-order interaction based on language context for sentence generation. In this way, we learn the object-relationship dependencies in different stages, which can provide abundant cues for both visual understanding and caption generation. Extensive experiments show that the proposed method can achieve competitive performances against the state-of-the-art methods on MSCOCO dataset. Additional ablation studies further validate its effectiveness.

Bidirectional interactive alignment network for image captioning

Bi-Directional Co-Attention Network for Image Captioning

BENet: bi-directional enhanced network for image captioning

HCNet: Hierarchical Feature Aggregation and Cross-Modal Feature Alignment for Remote Sensing Image Captioning

Center-enhanced video captioning model with multimodal semantic alignment

Dynamic-balanced Double-Attention Fusion for Image Captioning

Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching

Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning

CMGNet: Collaborative multi-modal graph network for video captioning

I3N: Intra- and Inter-representation Interaction Network for Change Captioning

Multimodal-enhanced hierarchical attention network for video captioning

Avtmnet: Adaptive Visual-Text Merging Network for Image Captioning

Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive Style

MAENet: A Novel Multi-Head Association Attention Enhancement Network for Completing Intra-Modal Interaction in Image Captioning

A Dual-Feature-Based Adaptive Shared Transformer Network for Image Captioning

iCAR: Bridging Image Classification and Image-text Alignment for Visual Recognition

A Cooperative Approach Based on Self-Attention with Interactive Attribute for Image Caption

BCMFIFuse: A Bilateral Cross-Modal Feature Interaction-Based Network for Infrared and Visible Image Fusion

Image Captioning with Deep Bidirectional LSTMs

High-Order Interaction Learning for Image Captioning

Cross-modal fine-grained alignment and fusion network for multimodal aspect-based sentiment analysis