Abstract:The fusion of far infrared (FIR) and visible images aims to generate a high-quality composite image that contains salient structures and abundant texture details for human visual perception. However, the existing fusion methods typically fall short of utilizing complementary source image characteristics to boost the features extracted from degraded visible or FIR images, thus they cannot generate satisfactory fusion results in adverse lighting or weather conditions. In this paper, we propose a novel Cross-Modal multispectral image Enhancement and Fusion framework (CMEFusion), which adaptively enhances both FIR and visible inputs by leveraging complementary cross-modal features to further facilitate multispectral feature aggregation. Specifically, we first present a new cross-modal image enhancement sub-network (CMIENet), which is built on a CNN-Transformer hybrid architecture to perform the complementary exchange of local-salient and global-contextual features extracted from FIR and visible modalities, respectively. Then, we design a gradient-content differential fusion sub-network (GCDFNet) to progressively integrate decoupled gradient and content information via modified central difference convolution. Finally, we present a comprehensive joint enhancement-fusion multi-term loss function to drive the model to narrow the optimization gap between the above-mentioned two sub-networks based on the self-supervised aspects of exposure, color, structure, and intensity. In this manner, the proposed CMEFusion model facilitates better-performing visible and FIR image fusion in an end-to-end way, achieving enhanced visual quality with more natural and realistic appearances. Extensive experiments validate that CMEFusion surpasses state-of-the-art image fusion algorithms, as evidenced by superior performance in both visual quality and quantitative evaluations.

Pctlfusion: A Progressive Fusion Network Via Contextual Texture Learning for Infrared and Visible Image

Infrared and Visible Image Fusion Based on a Two-Stage Class Conditioned Auto-Encoder Network.

PIAFusion: A progressive infrared and visible image fusion network based on illumination aware

[Reply to the discussion by Heinrich Kunze].

DCFusion: Dual-Headed Fusion Strategy and Contextual Information Awareness for Infrared and Visible Remote Sensing Image

TCCFusion: An Infrared and Visible Image Fusion Method based on Transformer and Cross Correlation

DCFusion: A Dual-Frequency Cross-Enhanced Fusion Network for Infrared and Visible Image Fusion.

Infrared and Visible Image Fusion Based on Filtering Enhancement

Adaptive low light visual enhancement and high-significant target detection for infrared and visible image fusion

Multi-scale attention-based lightweight network with dilated convolutions for infrared and visible image fusion

SFDFusion: An Efficient Spatial-Frequency Domain Fusion Network for Infrared and Visible Image Fusion

Infrared-visible Image Fusion Using Accelerated Convergent Convolutional Dictionary Learning

EV-Fusion: A Novel Infrared and Low-Light Color Visible Image Fusion Network Integrating Unsupervised Visible Image Enhancement

MIFFuse: A Multi-Level Feature Fusion Network for Infrared and Visible Images

Integrating Parallel Attention Mechanisms and Multi-Scale Features for Infrared and Visible Image Fusion

CMEFusion: Cross-Modal Enhancement and Fusion of FIR and Visible Images

Infrared and visible image fusion with entropy-based adaptive fusion module and mask-guided convolutional neural network

MAFusion: Multiscale Attention Network for Infrared and Visible Image Fusion

SPFusion: A multi-task semantic perception infrared and visible light fusion method with quality assessment

Visible and Infrared Image Fusion Based on Attention and Multiscale Residuals

TDDFusion: A Target-Driven Dual Branch Network for Infrared and Visible Image Fusion