Abstract:Infrared and visible image fusion aims to extract complementary features to synthesize a single fused image. Many methods employ convolutional neural networks (CNNs) to extract local features due to its translation invariance and locality. However, CNNs fail to consider the image's non-local self-similarity (NLss), though it can expand the receptive field by pooling operations, it still inevitably leads to information loss. In addition, the transformer structure extracts long-range dependence by considering the correlativity among all image patches, leading to information redundancy of such transformer-based methods. However, graph representation is more flexible than grid (CNN) or sequence (transformer structure) representation to address irregular objects, and graph can also construct the relationships among the spatially repeatable details or texture with far-space distance. Therefore, to address the above issues, it is significant to convert images into the graph space and thus adopt graph convolutional networks (GCNs) to extract NLss. This is because the graph can provide a fine structure to aggregate features and propagate information across the nearest vertices without introducing redundant information. Concretely, we implement a cascaded NLss extraction pattern to extract NLss of intra- and inter-modal by exploring interactions of different image pixels in intra- and inter-image positional distance. We commence by preforming GCNs on each intra-modal to aggregate features and propagate information to extract independent intra-modal NLss. Then, GCNs are performed on the concatenate intra-modal NLss features of infrared and visible images, which can explore the cross-domain NLss of inter-modal to reconstruct the fused image. Ablation studies and extensive experiments illustrates the effectiveness and superiority of the proposed method on three datasets.

CGTF: Convolution-Guided Transformer for Infrared and Visible Image Fusion

TCCFusion: An Infrared and Visible Image Fusion Method based on Transformer and Cross Correlation

THFuse: An Infrared and Visible Image Fusion Network using Transformer and Hybrid Feature Extractor

GTMFuse: Group-Attention Transformer-Driven Multiscale Dense Feature-Enhanced Network for Infrared and Visible Image Fusion

DTFusion: Infrared and Visible Image Fusion Based on Dense Residual PConv-ConvNeXt and Texture-Contrast Compensation

HDCTfusion: Hybrid Dual-Branch Network Based on CNN and Transformer for Infrared and Visible Image Fusion

A Deep Learning Framework for Infrared and Visible Image Fusion Without Strict Registration

SDTFusion: A split-head dense transformer based network for infrared and visible image fusion

CTFusion: CNN-transformer-based self-supervised learning for infrared and visible image fusion

Infrared and Visible Image Fusion Based on a Two-Stage Class Conditioned Auto-Encoder Network.

MFST: Multi-Modal Feature Self-Adaptive Transformer for Infrared and Visible Image Fusion

Infrared-visible Image Fusion Using Accelerated Convergent Convolutional Dictionary Learning

HDCCT: Hybrid Densely Connected CNN and Transformer for Infrared and Visible Image Fusion

HATF: Multi-Modal Feature Learning for Infrared and Visible Image Fusion via Hybrid Attention Transformer

HitFusion: Infrared and Visible Image Fusion for High-Level Vision Tasks Using Transformer

Graph Representation Learning for Infrared and Visible Image Fusion

Multi-scale attention-based lightweight network with dilated convolutions for infrared and visible image fusion

DATFuse: Infrared and Visible Image Fusion via Dual Attention Transformer

DCFusion: A Dual-Frequency Cross-Enhanced Fusion Network for Infrared and Visible Image Fusion.

YDTR: Infrared and Visible Image Fusion via Y-shape Dynamic Transformer

Infrared and Visible Image Fusion with Convolutional Neural Networks.