Abstract:Renal tumors are one of the common diseases of urology, and precise segmentation of these tumors plays a crucial role in aiding physicians to improve diagnostic accuracy and treatment effectiveness. Nevertheless, inherent challenges associated with renal tumors, such as indistinct boundaries, morphological variations, and uncertainties in size and location, segmenting renal tumors accurately remains a significant challenge in the field of medical image segmentation. With the development of deep learning, substantial achievements have been made in the domain of medical image segmentation. However, existing models lack specificity in extracting features of renal tumors across different network hierarchies, which results in insufficient extraction of renal tumor features and subsequently affects the accuracy of renal tumor segmentation. To address this issue, we propose the Selective Kernel, Vision Transformer, and Coordinate Attention Enhanced U-Net (STC-UNet). This model aims to enhance feature extraction, adapting to the distinctive characteristics of renal tumors across various network levels. Specifically, the Selective Kernel modules are introduced in the shallow layers of the U-Net, where detailed features are more abundant. By selectively employing convolutional kernels of different scales, the model enhances its capability to extract detailed features of renal tumors across multiple scales. Subsequently, in the deeper layers of the network, where feature maps are smaller yet contain rich semantic information, the Vision Transformer modules are integrated in a non-patch manner. These assist the model in capturing long-range contextual information globally. Their non-patch implementation facilitates the capture of fine-grained features, thereby achieving collaborative enhancement of global–local information and ultimately strengthening the model's extraction of semantic features of renal tumors. Finally, in the decoder segment, the Coordinate Attention modules embedding positional information are proposed aiming to enhance the model's feature recovery and tumor region localization capabilities. Our model is validated on the KiTS19 dataset, and experimental results indicate that compared to the baseline model, STC-UNet shows improvements of 1.60%, 2.02%, 2.27%, 1.18%, 1.52%, and 1.35% in IoU, Dice, Accuracy, Precision, Recall, and F1-score, respectively. Furthermore, the experimental results demonstrate that the proposed STC-UNet method surpasses other advanced algorithms in both visual effectiveness and objective evaluation metrics.

ST-Unet: Swin Transformer boosted U-Net with Cross-Layer Feature Enhancement for medical image segmentation

Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation

Swin-TransUper: Swin Transformer-based UperNet for medical image segmentation

STM-UNet: An Efficient U-shaped Architecture Based on Swin Transformer and Multi-scale MLP for Medical Image Segmentation

SAttisUNet: UNet-like Swin Transformer with Attentive Skip Connections for Enhanced Medical Image Segmentation

DS-TransUNet:Dual Swin Transformer U-Net for Medical Image Segmentation

Sfe-Transunet: A Transformer-Based U-Net With Skipped Features Enhancer For Medical Image Segmentation

DS-TransUNet: Dual Swin Transformer U-Net for Medical Image Segmentation

SSTrans-Net: Smart Swin Transformer Network for medical image segmentation

SW-UNet: a U-Net fusing sliding window transformer block with CNN for segmentation of lung nodules

CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation

A novel full-convolution UNet-transformer for medical image segmentation

FCTrans UNet: A Hybrid CNN and Transformer Model for Medical Image Segmentations

ConvWin-UNet: UNet-like hierarchical vision Transformer combined with convolution for medical image segmentation.

EG-TransUNet: a transformer-based U-Net with enhanced and guided models for biomedical image segmentation

TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation

Enhancing medical image segmentation with a multi-transformer U-Net

P-TransUNet: an improved parallel network for medical image segmentation

SWTRU: Star-shaped Window Transformer Reinforced U-Net for medical image segmentation

Swin-Net: A Swin-Transformer-Based Network Combing with Multi-Scale Features for Segmentation of Breast Tumor Ultrasound Images

STC-UNet: renal tumor segmentation based on enhanced feature extraction at different network levels