Abstract:The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM's segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM's mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM's mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner, facilitated by our effective robust training strategy (RTS). During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM's segmentation stability across a wide range of prompt qualities, while 2) retaining SAM's powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation (by 1 training epoch). Extensive experiments across multiple datasets validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes will be released upon acceptance. <a class="link-external link-https" href="https://github.com/fanq15/Stable-SAM" rel="external noopener nofollow">this https URL</a>

Understanding Segment Anything Model: SAM is Biased Towards Texture Rather than Shape

SAM Fails to Segment Anything? – SAM-Adapter: Adapting SAM in Underperformed Scenes: Camouflage, Shadow, Medical Image Segmentation, and More

SAMP: Adapting Segment Anything Model for Pose Estimation

Principles, applications, and advancements of the Segment Anything Model

Quantifying the Limits of Segment Anything Model: Analyzing Challenges in Segmenting Tree-Like and Low-Contrast Structures

A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering

BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model

Segment anything, from space?

Segment Anything with Multiple Modalities

Accuracy of Segment-Anything Model (SAM) in medical image segmentation tasks

There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks

Segment Anything Model (SAM) Meets Glass: Mirror and Transparent Objects Cannot Be Easily Detected

Semantic-SAM: Segment and Recognize Anything at Any Granularity

Segment Anything without Supervision

Stable Segment Anything Model

Promoting Segment Anything Model towards Highly Accurate Dichotomous Image Segmentation

Tuning a SAM-Based Model with Multi-Cognitive Visual Adapter to Remote Sensing Instance Segmentation

Computer-Vision Benchmark Segment-Anything Model (SAM) in Medical Images: Accuracy in 12 Datasets

Evaluating SAM2's Role in Camouflaged Object Detection: From SAM to SAM2

From SAM to SAM 2: Exploring Improvements in Meta's Segment Anything Model

Segment Anything Model for Medical Images?