Abstract:Medical image classification has developed rapidly under the impetus of the convolutional neural network (CNN). Due to the fixed size of the receptive field of the convolution kernel, it is difficult to capture the global features of medical images. Although the self-attention-based Transformer can model long-range dependencies, it has high computational complexity and lacks local inductive bias. Much research has demonstrated that global and local features are crucial for image classification. However, medical images have a lot of noisy, scattered features, intra-class variation, and inter-class similarities. This paper proposes a three-branch hierarchical multi-scale feature fusion network structure termed as HiFuse for medical image classification as a new method. It can fuse the advantages of Transformer and CNN from multi-scale hierarchies without destroying the respective modeling so as to improve the classification accuracy of various medical images. A parallel hierarchy of local and global feature blocks is designed to efficiently extract local features and global representations at various semantic scales, with the flexibility to model at different scales and linear computational complexity relevant to image size. Moreover, an adaptive hierarchical feature fusion block (HFF block) is designed to utilize the features obtained at different hierarchical levels comprehensively. The HFF block contains spatial attention, channel attention, residual inverted MLP, and shortcut to adaptively fuse semantic information between various scale features of each branch. The accuracy of our proposed model on the ISIC2018 dataset is 7.6% higher than baseline, 21.5% on the Covid-19 dataset, and 10.4% on the Kvasir dataset. Compared with other advanced models, the HiFuse model performs the best. Our code is open-source and available from <a class="link-external link-https" href="https://github.com/huoxiangzuo/HiFuse" rel="external noopener nofollow">this https URL</a>.

DeepFuse Neural Networks

ObjectFusion: an Object Detection and Segmentation Framework with RGB-D SLAM and Convolutional Neural Networks

ExFuse: Enhancing Feature Fusion for Semantic Segmentation

FAFuse: A Four-Axis Fusion framework of CNN and transformer for medical image segmentation

Aerial Scene Classification Via Multilevel Fusion Based on Deep Convolutional Neural Networks.

THFuse: An Infrared and Visible Image Fusion Network using Transformer and Hybrid Feature Extractor

A Late Fusion Approach for Harnessing Multi-Cnn Model High-Level Features

CoTrFuse: a novel framework by fusing CNN and transformer for medical image segmentation

TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation

DPACFuse: Dual-Branch Progressive Learning for Infrared and Visible Image Fusion with Complementary Self-Attention and Convolution

DenseFuse: A Fusion Approach to Infrared and Visible Images

Mask-R-FCN: A Deep Fusion Network for Semantic Segmentation.

Feature Fusion Use Unsupervised Prior Knowledge to Let Small Object Represent

Fused DNN: A deep neural network fusion approach to fast and robust pedestrian detection

GLFuse: A Global and Local Four-Branch Feature Extraction Network for Infrared and Visible Image Fusion

A Deep Fully Convolution Neural Network for Semantic Segmentation Based on Adaptive Feature Fusion

DM-Fusion: Deep Model-Driven Network for Heterogeneous Image Fusion.

CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion

Sub-pixel multi-scale fusion network for medical image segmentation

A multi-path adaptive fusion network for multimodal brain tumor segmentation

HiFuse: Hierarchical Multi-Scale Feature Fusion Network for Medical Image Classification