Abstract:Medical image classification has developed rapidly under the impetus of the convolutional neural network (CNN). Due to the fixed size of the receptive field of the convolution kernel, it is difficult to capture the global features of medical images. Although the self-attention-based Transformer can model long-range dependencies, it has high computational complexity and lacks local inductive bias. Much research has demonstrated that global and local features are crucial for image classification. However, medical images have a lot of noisy, scattered features, intra-class variation, and inter-class similarities. This paper proposes a three-branch hierarchical multi-scale feature fusion network structure termed as HiFuse for medical image classification as a new method. It can fuse the advantages of Transformer and CNN from multi-scale hierarchies without destroying the respective modeling so as to improve the classification accuracy of various medical images. A parallel hierarchy of local and global feature blocks is designed to efficiently extract local features and global representations at various semantic scales, with the flexibility to model at different scales and linear computational complexity relevant to image size. Moreover, an adaptive hierarchical feature fusion block (HFF block) is designed to utilize the features obtained at different hierarchical levels comprehensively. The HFF block contains spatial attention, channel attention, residual inverted MLP, and shortcut to adaptively fuse semantic information between various scale features of each branch. The accuracy of our proposed model on the ISIC2018 dataset is 7.6% higher than baseline, 21.5% on the Covid-19 dataset, and 10.4% on the Kvasir dataset. Compared with other advanced models, the HiFuse model performs the best. Our code is open-source and available from <a class="link-external link-https" href="https://github.com/huoxiangzuo/HiFuse" rel="external noopener nofollow">this https URL</a>.

CheXFusion: Effective Fusion of Multi-View Features using Transformers for Long-Tailed Chest X-Ray Classification

Bag of Tricks for Long-Tailed Multi-Label Classification on Chest X-Rays

TransFusion: Multi-view Divergent Fusion for Medical Image Segmentation with Transformers

CTransCNN: Combining transformer and CNN in multilabel medical image classification

MACTFusion: Lightweight Cross Transformer for Adaptive Multimodal Medical Image Fusion

Ensemble of ConvNeXt V2 and MaxViT for Long-Tailed CXR Classification with View-Based Aggregation

MultiFusionNet: Multilayer Multimodal Fusion of Deep Neural Networks for Chest X-Ray Image Classification

UniChest: Conquer-and-Divide Pre-training for Multi-Source Chest X-Ray Classification

Multi-Scale Feature Fusion using Parallel-Attention Block for COVID-19 Chest X-ray Diagnosis

LTCXNet: Advancing Chest X-Ray Analysis with Solutions for Long-Tailed Multi-Label Classification and Fairness Challenges

Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge

CheXDouble: Dual-Supervised interpretable disease diagnosis model

SynthEnsemble: A Fusion of CNN, Vision Transformer, and Hybrid Models for Multi-Label Chest X-Ray Classification

CTFC: A Convolution and Visual Transformer Based Classifier for Few-Shot Chest X-ray Images

FAFuse: A Four-Axis Fusion framework of CNN and transformer for medical image segmentation

MedFuse: Multi-modal fusion with clinical time-series data and chest X-ray images

MDC-RHT: Multi-Modal Medical Image Fusion via Multi-Dimensional Dynamic Convolution and Residual Hybrid Transformer

HiFuse: Hierarchical Multi-Scale Feature Fusion Network for Medical Image Classification

CFATransUnet: Channel-wise cross fusion attention and transformer for 2D medical image segmentation

CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion

Random Token Fusion for Multi-View Medical Diagnosis