Abstract:More diverse and closer to human-like captions are of paramount importance in image captioning. Recent research has achieved significant advancements, with the majority adopting end-to-end encoder-decoder architectures that integrate specific feature-text processing. However, the homogeneity of their model structures, the simplicity or complexity of featuretext fusion, and the uniformity of training objectives have all to some extent affected the diversity and effectiveness of caption generation, thus limiting the potential applications of this task. Therefore, in this paper, we propose the Regular Constrained Multimodal Fusion (RCMF) method for image captioning to better integrate information across and within modalities, while also approaching human-like fine-grained semantic perception and relationship reasoning capabilities. Initially, our RCMF preprocesses images using a Swin-Transformer and then an extended encoder with a new intra-modal fusion module, utilizing window-focused linear attention to capture features and leveraging refined grid and global visual features. By combining text features, RCMF employs a cross-modal fusion module and decoder to deeply model the interaction between text and image. Additionally, RCMF first introduces a new additional regulatory modal fusion reasoning (MFR) branch, which surpasses the above architectures. Its MFR loss combined with cross-entropy loss forms a new training objective strategy, effectively mining fine-grained relationships between images and text, perceiving the semantic information of images and their corresponding captions, thereby regulating the generated captions to be more diverse and human-like. Experimental results based on the MS COCO 2014 dataset, particularly under the same experimental conditions, demonstrate the outstanding performance of our method, especially in terms of METEOR, ROUGE-L, CIDEr, and SPICE metrics. Visualization results further intuitively confirm the effectiveness of our RCMF method. Source code in https://github.com/200084/RCMF-for-image-caption.

Cross-region Feature Fusion with Geometrical Relationship for OCR-based Image Captioning

Improving OCR-based Image Captioning by Incorporating Geometrical Relationship

Cross on Cross Attention: Deep Fusion Transformer for Image Captioning

Exploring Visual Relationships Via Transformer-based Graphs for Enhanced Image Captioning

Fine-Grained Features for Image Captioning

Exploring refined dual visual features cross-combination for image captioning

Improving Fusion of Region Features and Grid Features Via Two-Step Interaction for Image-Text Retrieval

Cooperative Connection Transformer for Remote Sensing Image Captioning

Exploring better image captioning with grid features

Dual-level Collaborative Transformer for Image Captioning

Generating Spatial-aware Captions for TextCaps

TSFNet: Triple-Steam Image Captioning

Scene captioning with deep fusion of images and point clouds

Exploring and Distilling Cross-Modal Information for Image Captioning

Dual visual align-cross attention-based image captioning transformer

Regular Constrained Multimodal Fusion for Image Captioning

Cross-scale Feature Fusion Self-attention for Image Captioning

Enhanced Transformer for Remote-Sensing Image Captioning with Positional-Channel Semantic Fusion

Region-guided Transformer for Remote Sensing Image Captioning

Image Captioning: Transforming Objects into Words

Feature Fusion Based on Transformer for Cross-modal Retrieval