Domain Generalization for Endoscopic Image Segmentation by Disentangling Style-Content Information and SuperPixel Consistency

Mansoor Ali Teevno,Rafael Martinez-Garcia-Pena,Gilberto Ochoa-Ruiz,Sharib Ali
DOI: https://doi.org/10.1109/CBMS61543.2024.00070
2024-09-19
Abstract:Frequent monitoring is necessary to stratify individuals based on their likelihood of developing gastrointestinal (GI) cancer precursors. In clinical practice, white-light imaging (WLI) and complementary modalities such as narrow-band imaging (NBI) and fluorescence imaging are used to assess risk areas. However, conventional deep learning (DL) models show degraded performance due to the domain gap when a model is trained on one modality and tested on a different one. In our earlier approach, we used a superpixel-based method referred to as "SUPRA" to effectively learn domain-invariant information using color and space distances to generate groups of pixels. One of the main limitations of this earlier work is that the aggregation does not exploit structural information, making it suboptimal for segmentation tasks, especially for polyps and heterogeneous color distributions. Therefore, in this work, we propose an approach for style-content disentanglement using instance normalization and instance selective whitening (ISW) for improved domain generalization when combined with SUPRA. We evaluate our approach on two datasets: EndoUDA Barrett's Esophagus and EndoUDA polyps, and compare its performance with three state-of-the-art (SOTA) methods. Our findings demonstrate a notable enhancement in performance compared to both baseline and SOTA methods across the target domain data. Specifically, our approach exhibited improvements of 14%, 10%, 8%, and 18% over the baseline and three SOTA methods on the polyp dataset. Additionally, it surpassed the second-best method (EndoUDA) on the Barrett's Esophagus dataset by nearly 2%.
Computer Vision and Pattern Recognition,Artificial Intelligence
What problem does this paper attempt to address?
The paper aims to address the issue of performance degradation in traditional deep learning models for endoscopic image segmentation due to domain gaps. Specifically, researchers have found that when a model is trained on one modality (such as White Light Imaging, WLI) but tested on another modality (such as Narrow Band Imaging, NBI), the model's performance significantly drops. To solve this problem, the authors propose a method that combines instance normalization and instance selective whitening (ISW) to decouple style and content information, thereby improving the model's generalization ability across different data domains. The main contribution of the paper is the proposal of an improved framework that maintains good segmentation performance across different endoscopic imaging modalities. Through experiments on two datasets (Barrett's Esophagus and Polyps), the new method achieves significant performance improvements on target domain data compared to baseline models and other state-of-the-art methods. Specifically, on the Polyps dataset, it improves by 14%, 10%, 8%, and 18% over the baseline and three other state-of-the-art methods, respectively. This indicates that the method can effectively suppress domain-specific information, allowing the model to learn more discriminative features from the source domain and perform better on the target domain.