Abstract:Zero-shot detection (ZSD) aims to locate and classify unseen objects in pictures or videos by semantic auxiliary information without additional training examples. Most of the existing ZSD methods are based on two-stage models, which achieve the detection of unseen classes by aligning object region proposals with semantic embeddings. However, these methods have several limitations, including poor region proposals for unseen classes, lack of consideration of semantic representations of unseen classes or their inter-class correlations, and domain bias towards seen classes, which can degrade overall performance. To address these issues, the Trans-ZSD framework is proposed, which is a transformer-based multi-scale contextual detection framework that explicitly exploits inter-class correlations between seen and unseen classes and optimizes feature distribution to learn discriminative features. Trans-ZSD is a single-stage approach that skips proposal generation and performs detection directly, allowing the encoding of long-term dependencies at multiple scales to learn contextual features while requiring fewer inductive biases. Trans-ZSD also introduces a foreground-background separation branch to alleviate the confusion of unseen classes and backgrounds, contrastive learning to learn inter-class uniqueness and reduce misclassification between similar classes, and explicit inter-class commonality learning to facilitate generalization between related classes. Trans-ZSD addresses the domain bias problem in end-to-end generalized zero-shot detection (GZSD) models by using balance loss to maximize response consistency between seen and unseen predictions, ensuring that the model does not bias towards seen classes. The Trans-ZSD framework is evaluated on the PASCAL VOC and MS COCO datasets, demonstrating significant improvements over existing ZSD models.

Temporal and cross-modal attention for audio-visual zero-shot learning

Audio-visual Generalised Zero-shot Learning with Cross-modal Attention and Language

Joint Learning of Attended Zero-Shot Features and Visual-Semantic Mapping.

Temporal–Semantic Aligning and Reasoning Transformer for Audio-Visual Zero-Shot Learning

Audio-visual Generalized Zero-shot Learning the Easy Way

Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models

Multi-label Zero-Shot Audio Classification with Temporal Attention

Coordinated Joint Multimodal Embeddings for Generalized Audio-Visual Zeroshot Classification and Retrieval of Videos

Multi-modal Generative Adversarial Network for Zero-Shot Learning

Transformer-Based Approach Via Contrastive Learning for Zero-Shot Detection.

Manifold Regularized Cross-Modal Embedding for Zero-Shot Learning

Multiscale Visual-Attribute Co-Attention for Zero-Shot Image Recognition

Learning Explicit and Implicit Latent Common Spaces for Audio-Visual Cross-Modal Retrieval

A Simple Framework for Open-Vocabulary Zero-Shot Segmentation

Multi-modal Multi-grained Embedding Learning for Generalized Zero-Shot Video Classification

Visual and Semantic Prototypes-Jointly Guided CNN for Generalized Zero-shot Learning

ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action Detection

ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning

Visual-guided attentive attributes embedding for zero-shot learning

An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment

A Discriminative Cross-Aligned Variational Autoencoder for Zero-Shot Learning