Abstract:Objective.Automatic mutli-organ segmentation from anotomical images is essential in disease diagnosis and treatment planning. The U-shaped neural network with encoder-decoder has achieved great success in various segmentation tasks. However, a pure convolutional neural network (CNN) is not suitable for modeling long-range relations due to limited receptive fields, and a pure transformer is not good at capturing pixel-level features.Approach.We propose a new hybrid network named MSCT-UNET which fuses CNN features with transformer features at multi-scale and introduces multi-task contrastive learning to improve the segmentation performance. Specifically, the multi-scale low-level features extracted from CNN are further encoded through several transformers to build hierarchical global contexts. Then the cross fusion block fuses the low-level and high-level features in different directions. The deep-fused features are flowed back to the CNN and transformer branch for the next scale fusion. We introduce multi-task contrastive learning including a self-supervised global contrast learning and a supervised local contrast learning into MSCT-UNET. We also make the decoder stronger by using a transformer to better restore the segmentation map.Results.Evaluation results on ACDC, Synapase and BraTS datasets demonstrate the improved performance over other methods compared. Ablation study results prove the effectiveness of our major innovations.Significance.The hybrid encoder of MSCT-UNET can capture multi-scale long-range dependencies and fine-grained detail features at the same time. The cross fusion block can fuse these features deeply. The multi-task contrastive learning of MSCT-UNET can strengthen the representation ability of the encoder and jointly optimize the networks. The source code is publicly available at:https://github.com/msctunet/MSCT_UNET.git.

Stitching Transformer and Convolution in U-Net for Medical Image Segmentation

Mixed Transformer U-Net for Medical Image Segmentation

Multi-scale Neighborhood Attention Transformer on U-Net for Medical Image Segmentation.

MSTCNet: Parallel Multi-Scale Network For Medical Image Segmentation.

TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation

UCTNet: Uncertainty-guided CNN-Transformer hybrid networks for medical image segmentation

MSCT-UNET: multi-scale contrastive transformer within U-shaped network for medical image segmentation

FCTrans UNet: A Hybrid CNN and Transformer Model for Medical Image Segmentations

U-Net Transformer: Self and Cross Attention for Medical Image Segmentation

HmsU-Net: A hybrid multi-scale U-net based on a CNN and transformer for medical image segmentation

Small Sample Image Segmentation by Coupling Convolutions and Transformers

TransU²-Net: An Effective Medical Image Segmentation Framework Based on Transformer and U²-Net

TSCA-Net: Transformer based spatial-channel attention segmentation network for medical images

A Multi-Scale Cross-Fusion Medical Image Segmentation Network Based on Dual-Attention Mechanism Transformer

A Improved Spatial-Reduction Attention Transformer for Medical Image Segmentation

Sfe-Transunet: A Transformer-Based U-Net With Skipped Features Enhancer For Medical Image Segmentation

CTC-Net: A Novel Coupled Feature-Enhanced Transformer and Inverted Convolution Network for Medical Image Segmentation

MTC-TransUNet: A Multi-Scale Mixed Convolution TransUNet for Medical Image Segmentation

3D TransUNet: Advancing Medical Image Segmentation through Vision Transformers

Trans-UNeter: A new Decoder of TransUNet for Medical Image Segmentation.

TransAttUnet: Multi-level Attention-guided U-Net with Transformer for Medical Image Segmentation.