Meta-Polyp: a baseline for efficient Polyp segmentation

Quoc-Huy Trinh
2023-07-18
Abstract:In recent years, polyp segmentation has gained significant importance, and many methods have been developed using CNN, Vision Transformer, and Transformer techniques to achieve competitive results. However, these methods often face difficulties when dealing with out-of-distribution datasets, missing boundaries, and small polyps. In 2022, Meta-Former was introduced as a new baseline for vision, which not only improved the performance of multi-task computer vision but also addressed the limitations of the Vision Transformer and CNN family backbones. To further enhance segmentation, we propose a fusion of Meta-Former with UNet, along with the introduction of a Multi-scale Upsampling block with a level-up combination in the decoder stage to enhance the texture, also we propose the Convformer block base on the idea of the Meta-former to enhance the crucial information of the local feature. These blocks enable the combination of global information, such as the overall shape of the polyp, with local information and boundary information, which is crucial for the decision of the medical segmentation. Our proposed approach achieved competitive performance and obtained the top result in the State of the Art on the CVC-300 dataset, Kvasir, and CVC-ColonDB dataset. Apart from Kvasir-SEG, others are out-of-distribution datasets. The implementation can be found at: <a class="link-external link-https" href="https://github.com/huyquoctrinh/MetaPolyp-CBMS2023" rel="external noopener nofollow">this https URL</a>.
Image and Video Processing,Computer Vision and Pattern Recognition,Machine Learning
What problem does this paper attempt to address?
The problem this paper attempts to address is improving the accuracy of polyp segmentation, particularly in handling out-of-distribution datasets, boundary deficiencies, and small polyps. Existing methods face difficulties in these areas; for example, Convolutional Neural Networks (CNNs) perform poorly in capturing global information, while Vision Transformers, although capable of capturing global information, lack in local information and boundary features. To solve these issues, the paper proposes a new method that combines MetaFormer with UNet and introduces multi-scale upsampling blocks and custom Convformer blocks to enhance texture and integrate local and global information. With these improvements, the paper aims to enhance the overall performance of polyp segmentation and achieve excellent results on multiple benchmark datasets.