Abstract:Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is learned from attribute descriptions shared between different classes, which act as strong priors for localizing object attributes that represent discriminative region features, enabling significant visual-semantic interaction. Although some attention-based models have attempted to learn such region features in a single image, the transferability and discriminative attribute localization of visual features are typically neglected. In this paper, we propose an attribute-guided Transformer network, termed TransZero, to refine visual features and learn attribute localization for discriminative visual embedding representations in ZSL. Specifically, TransZero takes a feature augmentation encoder to alleviate the cross-dataset bias between ImageNet and ZSL benchmarks, and improves the transferability of visual features by reducing the entangled relative geometry relationships among region features. To learn locality-augmented visual features, TransZero employs a visual-semantic decoder to localize the image regions most relevant to each attribute in a given image, under the guidance of semantic attribute information. Then, the locality-augmented visual features and semantic vectors are used to conduct effective visual-semantic interaction in a visual-semantic embedding network. Extensive experiments show that TransZero achieves the new state of the art on three ZSL benchmarks. The codes are available at: \url{<a class="link-external link-https" href="https://github.com/shiming-chen/TransZero" rel="external noopener nofollow">this https URL</a>}.

Discriminative and Robust Attribute Alignment for Zero-Shot Learning

Zero-shot Recognition with Latent Visual Attributes Learning.

Visual-guided attentive attributes embedding for zero-shot learning

Attribute self-representation steered by exclusive lasso for zero-shot learning

Learning complementary semantic information for zero-shot recognition

Learning discriminative visual semantic embedding for zero-shot recognition

Co-consistent Regularization with Discriminative Feature for Zero-Shot Learning

Learning Discriminative Projection with Visual Semantic Alignment for Generalized Zero Shot Learning.

Learning Discriminative Latent Attributes for Zero-Shot Classification.

Boosting Zero-shot Learning via Contrastive Optimization of Attribute Representations

Learning object-centric complementary features for zero-shot learning

Semantic-aware Visual Attributes Learning for Zero-Shot Recognition

Incorporating Attribute-Level Aligned Comparative Network for Generalized Zero-Shot Learning

Asymmetric Graph Based Zero Shot Learning

Towards Effective Deep Embedding for Zero-Shot Learning

Zero-shot Recognition with Image Attributes Generation Using Hierarchical Coupled Dictionary Learning

High-Discriminative Attribute Feature Learning for Generalized Zero-Shot Learning

Zero-Shot Learning via Discriminative Dual Semantic Auto-Encoder

Hybrid Regularization with Elastic Net and Linear Discriminant Analysis for Zero-Shot Image Recognition

TransZero: Attribute-guided Transformer for Zero-Shot Learning

Application of CLIP for Efficient Zero-Shot Learning