Abstract:Medical image classification has developed rapidly under the impetus of the convolutional neural network (CNN). Due to the fixed size of the receptive field of the convolution kernel, it is difficult to capture the global features of medical images. Although the self-attention-based Transformer can model long-range dependencies, it has high computational complexity and lacks local inductive bias. Much research has demonstrated that global and local features are crucial for image classification. However, medical images have a lot of noisy, scattered features, intra-class variation, and inter-class similarities. This paper proposes a three-branch hierarchical multi-scale feature fusion network structure termed as HiFuse for medical image classification as a new method. It can fuse the advantages of Transformer and CNN from multi-scale hierarchies without destroying the respective modeling so as to improve the classification accuracy of various medical images. A parallel hierarchy of local and global feature blocks is designed to efficiently extract local features and global representations at various semantic scales, with the flexibility to model at different scales and linear computational complexity relevant to image size. Moreover, an adaptive hierarchical feature fusion block (HFF block) is designed to utilize the features obtained at different hierarchical levels comprehensively. The HFF block contains spatial attention, channel attention, residual inverted MLP, and shortcut to adaptively fuse semantic information between various scale features of each branch. The accuracy of our proposed model on the ISIC2018 dataset is 7.6% higher than baseline, 21.5% on the Covid-19 dataset, and 10.4% on the Kvasir dataset. Compared with other advanced models, the HiFuse model performs the best. Our code is open-source and available from <a class="link-external link-https" href="https://github.com/huoxiangzuo/HiFuse" rel="external noopener nofollow">this https URL</a>.

What problem does this paper attempt to address?

### Problems the Paper Attempts to Solve This paper aims to address several key issues in medical image classification: 1. **Difficulty in Capturing Global Features**: Due to the fixed receptive field of convolutional kernels, traditional Convolutional Neural Networks (CNNs) find it challenging to capture the global features of medical images. 2. **Lack of Local Feature Details**: Although Transformers based on self-attention mechanisms can model long-range dependencies, their high computational complexity and lack of local inductive bias result in insufficient local feature details. 3. **Intra-class Variation and Inter-class Similarity**: Medical images have a large amount of noise, scattered features, intra-class variation, and inter-class similarity, making it difficult for models to distinguish between different categories of images. 4. **Fusion of Multi-scale Features**: To improve classification accuracy, it is necessary to effectively fuse local and global features at different levels. To this end, the paper proposes a new three-branch hierarchical multi-scale feature fusion network structure, called HiFuse, for medical image classification. HiFuse combines the advantages of Transformers and CNNs, efficiently extracting local features and global representations through a multi-scale hierarchical structure, and designs an adaptive hierarchical feature fusion block (HFF block) to fully utilize feature information at different levels. This structure not only improves classification accuracy but also reduces computational complexity.

HiFuse: Hierarchical Multi-Scale Feature Fusion Network for Medical Image Classification

FAFuse: A Four-Axis Fusion framework of CNN and transformer for medical image segmentation

Sub-pixel multi-scale fusion network for medical image segmentation

Multi-scale Feature Fusion Convolutional Neural Network for Multi-Modal Medical Image Fusion.

Multi-modal medical image fusion based on densely-connected high-resolution CNN and hybrid transformer

CMFuse: Correlation-based Multi-Scale Feature Fusion Network for the Detection of COVID-19 from Chest X-ray Images

An Improved Hybrid Network With a Transformer Module for Medical Image Fusion

A Local-Global Attention Fusion Framework with Tensor Decomposition for Medical Diagnosis

TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation

ECSFF: Exploring Efficient Cross-Scale Feature Fusion for Medical Image Segmentation.

CoTrFuse: a novel framework by fusing CNN and transformer for medical image segmentation

MFH‐Net: A Hybrid CNN‐Transformer Network Based Multi‐Scale Fusion for Medical Image Segmentation

Multi-Scale Fusion Global Feature Extraction Network for Multi-Modal Medical Image Fusion

MDC-RHT: Multi-Modal Medical Image Fusion via Multi-Dimensional Dynamic Convolution and Residual Hybrid Transformer

Microscopic Hyperspectral Image Classification Based on Fusion Transformer with Parallel CNN

Improving High Resolution Histology Image Classification with Deep Spatial Fusion Network

A Medical Image Fusion Method Based on Convolutional Neural Networks

Hahn-PCNN-CNN: an end-to-end multi-modal brain medical image fusion framework useful for clinical diagnosis

CASF-Net: Cross-attention and Cross-scale Fusion Network for Medical Image Segmentation

Multi-Modal Medical Image Fusion Based on FusionNet in YIQ Color Space

Multi-Scale Convolution-Transformer Fusion Network for Endoscopic Image Segmentation.