Abstract:BACKGROUND:Multi-modal learning is widely adopted to learn the latent complementary information between different modalities in multi-modal medical image segmentation tasks. Nevertheless, the traditional multi-modal learning methods require spatially well-aligned and paired multi-modal images for supervised training, which cannot leverage unpaired multi-modal images with spatial misalignment and modality discrepancy. For training accurate multi-modal segmentation networks using easily accessible and low-cost unpaired multi-modal images in clinical practice, unpaired multi-modal learning has received comprehensive attention recently.PURPOSE:Existing unpaired multi-modal learning methods usually focus on the intensity distribution gap but ignore the scale variation problem between different modalities. Besides, within existing methods, shared convolutional kernels are frequently employed to capture common patterns in all modalities, but they are typically inefficient at learning global contextual information. On the other hand, existing methods highly rely on a large number of labeled unpaired multi-modal scans for training, which ignores the practical scenario when labeled data is limited. To solve the above problems, we propose a modality-collaborative convolution and transformer hybrid network (MCTHNet) using semi-supervised learning for unpaired multi-modal segmentation with limited annotations, which not only collaboratively learns modality-specific and modality-invariant representations, but also could automatically leverage extensive unlabeled scans for improving performance.METHODS:We make three main contributions to the proposed method. First, to alleviate the intensity distribution gap and scale variation problems across modalities, we develop a modality-specific scale-aware convolution (MSSC) module that can adaptively adjust the receptive field sizes and feature normalization parameters according to the input. Secondly, we propose a modality-invariant vision transformer (MIViT) module as the shared bottleneck layer for all modalities, which implicitly incorporates convolution-like local operations with the global processing of transformers for learning generalizable modality-invariant representations. Third, we design a multi-modal cross pseudo supervision (MCPS) method for semi-supervised learning, which enforces the consistency between the pseudo segmentation maps generated by two perturbed networks to acquire abundant annotation information from unlabeled unpaired multi-modal scans.RESULTS:Extensive experiments are performed on two unpaired CT and MR segmentation datasets, including a cardiac substructure dataset derived from the MMWHS-2017 dataset and an abdominal multi-organ dataset consisting of the BTCV and CHAOS datasets. Experiment results show that our proposed method significantly outperforms other existing state-of-the-art methods under various labeling ratios, and achieves a comparable segmentation performance close to single-modal methods with fully labeled data by only leveraging a small portion of labeled data. Specifically, when the labeling ratio is 25%, our proposed method achieves overall mean DSC values of 78.56% and 76.18% in cardiac and abdominal segmentation, respectively, which significantly improves the average DSC value of two tasks by 12.84% compared to single-modal U-Net models.CONCLUSIONS:Our proposed method is beneficial for reducing the annotation burden of unpaired multi-modal medical images in clinical applications.

Multi-scale Transformer Network with Edge-aware Pre-training for Cross-Modality MR Image Synthesis

SLMT-Net: A Self-supervised Learning Based Multi-scale Transformer Network for Cross-Modality MR Image Synthesis

Rethinking Multi-Contrast MRI Super-Resolution: Rectangle-Window Cross-Attention Transformer and Arbitrary-Scale Upsampling

A Novel Method Of Synthetic Ct Generation From Mr Images Based On Convolutional Neural Networks

Mixed Transformer U-Net for Medical Image Segmentation

Multi-scale Tokens-Aware Transformer Network for Multi-region and Multi-sequence MR-to-CT Synthesis in A Single Model

A layer-wise fusion network incorporating self-supervised learning for multimodal MR image synthesis

Hi-Net: Hybrid-fusion Network for Multi-modal MR Image Synthesis

Enhancing CT Image synthesis from multi-modal MRI data based on a multi-task neural network framework

Multimodal MR Image Synthesis Using Gradient Prior and Adversarial Learning

Multi-Modality MR Image Synthesis via Confidence-Guided Aggregation and Cross-Modality Refinement

Synthesizing Multi-Contrast MR Images Via Novel 3D Conditional Variational Auto-Encoding GAN

Cross Attention Multi Scale CNN-Transformer Hybrid Encoder is General Medical Image Learner.

Multi-modal Modality-masked Diffusion Network for Brain MRI Synthesis with Random Modality Missing

Disentangled Multimodal Brain MR Image Translation via Transformer-based Modality Infuser

Learning Unified Hyper-Network for Multi-Modal MR Image Synthesis and Tumor Segmentation With Missing Modalities

Multi-Modal Transformer for Accelerated MR Imaging

Multimodal Transformer for Accelerated MR Imaging

Multi-modality MRI fusion with patch complementary pre-training for internet of medical things-based smart healthcare

A Modality-Collaborative Convolution and Transformer Hybrid Network for Unpaired Multi-Modal Medical Image Segmentation with Limited Annotations

Coarse-to-Fine Learning Framework for Semi-supervised Multimodal MRI Synthesis