Abstract:The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM's segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM's mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM's mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner, facilitated by our effective robust training strategy (RTS). During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM's segmentation stability across a wide range of prompt qualities, while 2) retaining SAM's powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation (by 1 training epoch). Extensive experiments across multiple datasets validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes will be released upon acceptance. <a class="link-external link-https" href="https://github.com/fanq15/Stable-SAM" rel="external noopener nofollow">this https URL</a>

SAMControl: Controlling Pose and Object for Image Editing with Soft Attention Mask

AM-SAM: Automated Prompting and Mask Calibration for Segment Anything Model

FocalClick: Towards Practical Interactive Image Segmentation.

FocSAM: Delving Deeply into Focused Objects in Segmenting Anything

BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model

SAMP: Adapting Segment Anything Model for Pose Estimation

SAM-REF: Rethinking Image-Prompt Synergy for Refinement in Segment Anything

Stable Segment Anything Model

Enhancing Text-to-Image Editing via Hybrid Mask-Informed Fusion

Masked-Attention Diffusion Guidance for Spatially Controlling Text-to-Image Generation

InstructEdit: Improving Automatic Masks for Diffusion-based Image Editing With User Instructions

Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing

MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted Guidance

Text-Guided Mask-free Local Image Retouching

PA-SAM: Prompt Adapter SAM for High-Quality Image Segmentation

AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image Editing

Variance-insensitive and Target-preserving Mask Refinement for Interactive Image Segmentation

DesignEdit: Multi-Layered Latent Decomposition and Fusion for Unified & Accurate Image Editing

Self-Attention-Masking Semantic Decomposition and Segmentation for Facial Attribute Manipulation

Tuning-Free Inversion-Enhanced Control for Consistent Image Editing

Segment Any Object Model (SAOM): Real-to-Simulation Fine-Tuning Strategy for Multi-Class Multi-Instance Segmentation