Abstract:Zero-shot detection (ZSD) aims to locate and classify unseen objects in pictures or videos by semantic auxiliary information without additional training examples. Most of the existing ZSD methods are based on two-stage models, which achieve the detection of unseen classes by aligning object region proposals with semantic embeddings. However, these methods have several limitations, including poor region proposals for unseen classes, lack of consideration of semantic representations of unseen classes or their inter-class correlations, and domain bias towards seen classes, which can degrade overall performance. To address these issues, the Trans-ZSD framework is proposed, which is a transformer-based multi-scale contextual detection framework that explicitly exploits inter-class correlations between seen and unseen classes and optimizes feature distribution to learn discriminative features. Trans-ZSD is a single-stage approach that skips proposal generation and performs detection directly, allowing the encoding of long-term dependencies at multiple scales to learn contextual features while requiring fewer inductive biases. Trans-ZSD also introduces a foreground-background separation branch to alleviate the confusion of unseen classes and backgrounds, contrastive learning to learn inter-class uniqueness and reduce misclassification between similar classes, and explicit inter-class commonality learning to facilitate generalization between related classes. Trans-ZSD addresses the domain bias problem in end-to-end generalized zero-shot detection (GZSD) models by using balance loss to maximize response consistency between seen and unseen predictions, ensuring that the model does not bias towards seen classes. The Trans-ZSD framework is evaluated on the PASCAL VOC and MS COCO datasets, demonstrating significant improvements over existing ZSD models.

Frustratingly Simple but Effective Zero-shot Detection and Segmentation: Analysis and a Strong Baseline

Zero-Shot Detection with Transferable Object Proposal Mechanism.

Weakly Supervised Classification Model for Zero‐shot Semantic Segmentation

A Simple Baseline for Zero-shot Semantic Segmentation with Pre-trained Vision-language Model.

Transformer-Based Approach Via Contrastive Learning for Zero-Shot Detection.

A Survey of Zero Shot Detection: Methods and Applications

A Survey of Deep Learning for Low-Shot Object Detection

Zero-Shot Co-salient Object Detection Framework

Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion

Zero-Shot Object Detection by Semantics-Aware DETR with Adaptive Contrastive Loss.

Zero-Shot Object Detection by Hybrid Region Embedding

Delving into Shape-aware Zero-shot Semantic Segmentation

Semantic-Promoted Debiasing and Background Disambiguation for Zero-Shot Instance Segmentation

See More and Know More: Zero-shot Point Cloud Segmentation via Multi-modal Visual Data

Visual semantic segmentation based on few/zero-shot learning: An overview

Zero-shot Unsupervised Transfer Instance Segmentation

Objectness-Aware Few-Shot Semantic Segmentation

Robust Region Feature Synthesizer for Zero-Shot Object Detection

A Simple Framework for Open-Vocabulary Zero-Shot Segmentation

AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation