Abstract:Few-shot Semantic Segmentation (FSS) attempts to segment the new category with only a few labeled samples, presenting a significant challenge. Existing approaches primarily focus on leveraging category information from the support set to identify objects of the new category in the query image. However, these models often struggle when confronted with substantial differences between paired images. To address issues stemming from scenario differences and intra-class diversity, this paper proposes an adaptive similarity-guided self-merging network. Firstly, style differences of multi-level features are introduced to alleviate the network's sensitivity to scenario variations and learn an adaptive weight for the K-shot scheme. Secondly, a feature-mask bi-aggregation module is designed to learn an enhanced feature and an initial mask for the query image. Within this module, dynamic correlations cover all the spatial locations, providing global information crucial for feature and mask aggregation. Subsequently, a self-merging module is proposed to alleviate prototype bias. It merges a self-prototype derived from the initial mask with an adaptive weighted support prototype obtained from K support images. Finally, the target object is segmented using the enhanced feature and merging prototype, and segmentation results are further refined by predictions of base categories and an adjustment factor derived from multilevel style differences. The proposed method achieves 69.1% (1-shot) and 72.3% (5-shot) mIoU on the PASCAL-5i dataset, and 47.4% (1-shot) and 52.1% (5-shot) mIoU on the COCO-20i dataset. These results demonstrate state-of-the-art segmentation performance compared to mainstream methods. (c) 2017 Elsevier Inc. All rights reserved.

CFENet: Leveraging CLIP Text Features for Enhanced Few-Shot Semantic Segmentation

Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation

CGMGM: A Cross-Gaussian Mixture Generative Model for Few-Shot Semantic Segmentation

Adaptive FSS: A Novel Few-Shot Segmentation Framework Via Prototype Enhancement

Iterative Few-shot Semantic Segmentation from Image Label Text

Memory-guided Network with Uncertainty-based Feature Augmentation for Few-shot Semantic Segmentation

Multi-Similarity Enhancement Network for Few-Shot Segmentation.

FFNet: Feature Fusion Network for Few-shot Semantic Segmentation

Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and Beyond

Masked Cross-image Encoding for Few-shot Segmentation

Few-Shot Segmentation via Channel Attention and Supervision Augmentation

Triple Attention Feature Aggregation for Few-Shot Segmentation

AFANet: Adaptive Frequency-Aware Network for Weakly-Supervised Few-Shot Semantic Segmentation

CLIP-Driven Prototype Network for Few-Shot Semantic Segmentation

Few-Shot Semantic Segmentation with Cyclic Memory Network

On filling the intra-class and inter-class gaps for few-shot segmentation

Multimodality Helps Few-Shot 3D Point Cloud Semantic Segmentation

Spatial Correlation Fusion Network for Few-Shot Segmentation

Mining Latent Classes for Few-shot Segmentation

Adaptive Similarity-Guided Self-Merging Network for Few-Shot Semantic Segmentation

MCEENet: Multi-Scale Context Enhancement and Edge-Assisted Network for Few-Shot Semantic Segmentation