2nd Place Winning Solution for the CVPR2023 Visual Anomaly and Novelty Detection Challenge: Multimodal Prompting for Data-centric Anomaly Detection

Yunkang Cao,Xiaohao Xu,Chen Sun,Yuqi Cheng,Liang Gao,Weiming Shen

2023-09-05

Abstract:This technical report introduces the winning solution of the team Segment Any Anomaly for the CVPR2023 Visual Anomaly and Novelty Detection (VAND) challenge. Going beyond uni-modal prompt, e.g., language prompt, we present a novel framework, i.e., Segment Any Anomaly + (SAA$+$), for zero-shot anomaly segmentation with multi-modal prompts for the regularization of cascaded modern foundation models. Inspired by the great zero-shot generalization ability of foundation models like Segment Anything, we first explore their assembly (SAA) to leverage diverse multi-modal prior knowledge for anomaly localization. Subsequently, we further introduce multimodal prompts (SAA$+$) derived from domain expert knowledge and target image context to enable the non-parameter adaptation of foundation models to anomaly segmentation. The proposed SAA$+$ model achieves state-of-the-art performance on several anomaly segmentation benchmarks, including VisA and MVTec-AD, in the zero-shot setting. We will release the code of our winning solution for the CVPR2023 VAN.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The paper aims to address the problem of Zero-shot Anomaly Segmentation (ZSAS), which involves segmenting various anomalies in images without using any normal or abnormal samples. Specifically, the research team proposed a new framework called "Segment Any Anomaly +" (SAA+), which enhances the capability of the base model through multimodal prompts, thereby improving the accuracy of anomaly detection on unknown objects. The core contributions of the paper are: 1. **Addressing the limitations of single-language prompts**: Traditional single-language prompts (e.g., "anomaly") are prone to false positives due to the domain gap between pre-trained data and downstream tasks. Therefore, SAA+ introduces multimodal prompts derived from domain expert knowledge and the context of the target image to reduce language ambiguity. 2. **Regularization using multimodal prompts**: By combining domain expert knowledge (e.g., descriptions of specific types of anomalies) and contextual information from the target image (e.g., visual saliency and confidence ranking), the base model's ability to identify anomalous regions is improved. 3. **Outstanding experimental results**: The SAA+ method achieved significantly better performance than other existing methods on multiple benchmark datasets (including VisA and MVTec-AD), demonstrating its effectiveness in zero-shot anomaly segmentation tasks.

2nd Place Winning Solution for the CVPR2023 Visual Anomaly and Novelty Detection Challenge: Multimodal Prompting for Data-centric Anomaly Detection

Segment Any Anomaly without Training via Hybrid Prompt Regularization

Adapting the Segment Anything Model for Multi-modal Retinal Anomaly Detection and Localization

VL4AD: Vision-Language Models Improve Pixel-wise Anomaly Detection

APRIL-GAN: A Zero-/Few-Shot Anomaly Classification and Segmentation Method for CVPR 2023 VAND Workshop Challenge Tracks 1&2: 1st Place on Zero-shot AD and 4th Place on Few-shot AD

Zero-Shot Anomaly Detection with Pre-trained Segmentation Models

Human-Free Automated Prompting for Vision-Language Anomaly Detection: Prompt Optimization with Meta-guiding Prompt Scheme

Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning

Learn Suspected Anomalies from Event Prompts for Video Anomaly Detection

Fine-grained Abnormality Prompt Learning for Zero-shot Anomaly Detection

Context-aware Feature Reconstruction for Class-Incremental Anomaly Detection and Localization

Anomaly Detection by Adapting a pre-trained Vision Language Model

AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis

PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection

A Diffusion-Based Framework for Multi-Class Anomaly Detection

Unsupervised Continual Anomaly Detection with Contrastively-learned Prompt

Learning Prompt-Enhanced Context Features for Weakly-Supervised Video Anomaly Detection

AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection

Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning

VCP-CLIP: A visual context prompting model for zero-shot anomaly segmentation

Anomaly-Aware Semantic Segmentation by Leveraging Synthetic-Unknown Data