Abstract:Background and Objective: Transformer, which is notable for its ability of global context modeling, has been used to remedy the shortcomings of Convolutional neural networks (CNN) and break its dominance in medical image segmentation. However, the self-attention module is both memory and computational inefficient, so many methods have to build their Transformer branch upon largely downsampled feature maps or adopt the tokenized image patches to fit their model into accessible GPUs. This patch-wise operation restricts the network in extracting pixel-level intrinsic structural or dependencies inside each patch, hurting the performance of pixel-level classification tasks. Methods: To tackle these issues, we propose a memory- and computation-efficient self-attention module to enable reasoning on relatively high-resolution features, promoting the efficiency of learning global information while effective grasping fine spatial details. Furthermore, we design a novel Multi-Branch Transformer (MultiTrans) architecture to provide hierarchical features for handling objects with variable shapes and sizes in medical images. By building four parallel Transformer branches on different levels of CNN, our hybrid network aggregates both multi-scale global contexts and multi-scale local features. Results: MultiTrans achieves the highest segmentation accuracy on three medical image datasets with different modalities: Synapse, ACDC and M&Ms. Compared to the Standard Self-Attention (SSA), the proposed Efficient Self-Attention (ESA) can largely reduce the training memory and computational complexity while even slightly improve the accuracy. Specifically, the training memory cost, FLOPs and Params of our ESA are 18.77%, 20.68% and 74.07% of the SSA. Conclusions: Experiments on three medical image datasets demonstrate the generality and robustness of the designed network. The ablation study shows the efficiency and effectiveness of our proposed ESA. Code is available at: https://github.com/Yanhua-Zhang/MultiTrans-extension .

SEMI-CONTRANS: Semi-Supervised Medical Image Segmentation via Multi-Scale Feature Fusion and Cross Teaching of CNN and Transformer

MedFCT: A Frequency Domain Joint CNN-Transformer Network for Semi-supervised Medical Image Segmentation

Semi-Supervised Convolutional Vision Transformer with Bi-Level Uncertainty Estimation for Medical Image Segmentation

Combinatorial CNN-Transformer Learning with Manifold Constraints for Semi-supervised Medical Image Segmentation

SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic Segmentation

MixFormer: a Mixed CNN-Transformer Backbone for Medical Image Segmentation

Context-aware and local-aware fusion with transformer for medical image segmentation

Lagrange Duality and Compound Multi-Attention Transformer for Semi-Supervised Medical Image Segmentation

FTransCNN: Fusing Transformer and a CNN based on fuzzy logic for uncertain medical image segmentation

TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation

Semi-Supervised Medical Image Segmentation Based on Deep Consistent Collaborative Learning

Multi-dimensional Fusion and Consistency for Semi-supervised Medical Image Segmentation

MultiTrans: Multi-branch transformer network for medical image segmentation

MSCT-UNET: multi-scale contrastive transformer within U-shaped network for medical image segmentation

Sub-pixel multi-scale fusion network for medical image segmentation

Dual encoder network with transformer-CNN for multi-organ segmentation

Dual CNN cross-teaching semi-supervised segmentation network with multi-kernels and global contrastive loss in ACDC

CFATransUnet: Channel-wise cross fusion attention and transformer for 2D medical image segmentation

CASF-Net: Cross-attention and Cross-scale Fusion Network for Medical Image Segmentation

When CNN Meet with ViT: Towards Semi-Supervised Learning for Multi-Class Medical Image Semantic Segmentation

CiT-Net: Convolutional Neural Networks Hand in Hand with Vision Transformers for Medical Image Segmentation