Abstract:Zero-shot detection (ZSD) aims to locate and classify unseen objects in pictures or videos by semantic auxiliary information without additional training examples. Most of the existing ZSD methods are based on two-stage models, which achieve the detection of unseen classes by aligning object region proposals with semantic embeddings. However, these methods have several limitations, including poor region proposals for unseen classes, lack of consideration of semantic representations of unseen classes or their inter-class correlations, and domain bias towards seen classes, which can degrade overall performance. To address these issues, the Trans-ZSD framework is proposed, which is a transformer-based multi-scale contextual detection framework that explicitly exploits inter-class correlations between seen and unseen classes and optimizes feature distribution to learn discriminative features. Trans-ZSD is a single-stage approach that skips proposal generation and performs detection directly, allowing the encoding of long-term dependencies at multiple scales to learn contextual features while requiring fewer inductive biases. Trans-ZSD also introduces a foreground-background separation branch to alleviate the confusion of unseen classes and backgrounds, contrastive learning to learn inter-class uniqueness and reduce misclassification between similar classes, and explicit inter-class commonality learning to facilitate generalization between related classes. Trans-ZSD addresses the domain bias problem in end-to-end generalized zero-shot detection (GZSD) models by using balance loss to maximize response consistency between seen and unseen predictions, ensuring that the model does not bias towards seen classes. The Trans-ZSD framework is evaluated on the PASCAL VOC and MS COCO datasets, demonstrating significant improvements over existing ZSD models.

Hierarchical contrastive representation for zero shot learning

Cluster-based Contrastive Disentangling for Generalized Zero-Shot Learning

Contrastive Visual Feature Filtering for Generalized Zero-Shot Learning

HCSC: Hierarchical Contrastive Selective Coding

Transformer-Based Approach Via Contrastive Learning for Zero-Shot Detection.

Hyperbolic Contrastive Learning for Visual Representations beyond Objects

Hyperspherical Consistency Regularization.

Hierarchically Decoupled Spatial-Temporal Contrast for Self-supervised Video Representation Learning

Unbiased Hybrid Generation Network for Zero-Shot Learning

HierarchicalContrast: A Coarse-to-Fine Contrastive Learning Framework for Cross-Domain Zero-Shot Slot Filling

Boosting Zero-shot Learning via Contrastive Optimization of Attribute Representations

Unsupervised Graph-Level Representation Learning with Hierarchical Contrasts

Multi-granularity contrastive zero-shot learning model based on attribute decomposition

Use All The Labels: A Hierarchical Multi-Label Contrastive Learning Framework

Compositional Zero-Shot Learning with Contextualized Cues and Adaptive Contrastive Training

Heterogeneous Contrastive Learning: Encoding Spatial Information for Compact Visual Representations

Semantics-Guided Contrastive Network for Zero-Shot Object detection

HAPiCLR: heuristic attention pixel-level contrastive loss representation learning for self-supervised pretraining

Learning Joint Feature Adaptation for Zero-Shot Recognition

CoHOZ: Contrastive Multimodal Prompt Tuning for Hierarchical Open-set Zero-shot Recognition

DUET: Cross-modal Semantic Grounding for Contrastive Zero-shot Learning