Abstract:Background and Objectives: Transformer is a model relying entirely on self-attention which has a wide range of applications in the field of natural language processing. Researchers are beginning to focus on the transformer in medical images due to the past few years having seen the rapid development of transformer in many vision fields such as vision transformer (ViT) and Swin transformer. In the last year, moreover, many scholars have applied transformer to medical image segmentation and have achieved good segmentation results. Transformer-based medical image segmentation has become one of the hot spots in this field. The purpose of this work is to categorize and review the segmentation methods of Unet-based transformer and other model based transformer in medical images. Methods: This paper summarizes the transformer-based segmentation models in the abdominal organs, heart, brain, and lung based on the relevant studies in the last two years. We described and analyzed the model structure including the position of the transformer in the model, the changes made by scholars to transformer and the combination with the model. In this work, the segmentation performance results based on Dice evaluation metrics are compared. Results: Through the help of 93 references, we find that researchers prefer to use Unet-based transformer models than others and place the transformer structure in the encoder. These new models improve the segmentation performance compared with U-Net and other segmentation models. However, there are not many related studies on lungs, which points to a new way for future research. Conclusions: We found that the combination of U-Net and transformer is more suitable for segmentation. In future research on medical image segmentation, researchers can use a suitable transformer-based segmentation method or modify the transformer structure according to the segmentation requirements. We hope that this work will be helpful for improvements of the transformer to solve clinical problems in medicine.

U-Net Transformer: Self and Cross Attention for Medical Image Segmentation

Mixed Transformer U-Net for Medical Image Segmentation

UTNet: A Hybrid Transformer Architecture for Medical Image Segmentation

Contextual Attention Network: Transformer Meets U-Net

TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation

Multi-scale Neighborhood Attention Transformer on U-Net for Medical Image Segmentation.

A Multi-Scale Cross-Fusion Medical Image Segmentation Network Based on Dual-Attention Mechanism Transformer

3D TransUNet: Advancing Medical Image Segmentation through Vision Transformers

TSCA-Net: Transformer based spatial-channel attention segmentation network for medical images

A Improved Spatial-Reduction Attention Transformer for Medical Image Segmentation

TransAttUnet: Multi-level Attention-guided U-Net with Transformer for Medical Image Segmentation.

Stitching Transformer and Convolution in U-Net for Medical Image Segmentation

U-Netmer: U-Net meets Transformer for medical image segmentation

UCTNet: Uncertainty-guided CNN-Transformer hybrid networks for medical image segmentation

TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers

MSCT-UNET: multi-scale contrastive transformer within U-shaped network for medical image segmentation

A novel full-convolution UNet-transformer for medical image segmentation

Sfe-Transunet: A Transformer-Based U-Net With Skipped Features Enhancer For Medical Image Segmentation

DA-TransUNet: Integrating Spatial and Channel Dual Attention with Transformer U-Net for Medical Image Segmentation

Transformer-based heart organ segmentation using a novel axial attention and fusion mechanism

Transformers in medical image segmentation: A review