Abstract:Compositional Zero-Shot Learning (CZSL) is a particular Zero-Shot Learning (ZSL) task that aims to utilize known concepts (e.g., states and objects) to identify novel state-object compositions for Image Classification. Previous works have primarily focused on disentangling concept compositions or exploring the complex interactions between the states and objects while neglecting the critical fact that the inference of many states and compositions is related to different frequency components, which should be analyzed from the global perspective. Therefore, we propose a Spatial-frequency Feature Fusion Network (SFFNet) to introduce a new branch that utilizes a frequency-domain filtering encoder to enhance key frequency components and capture non-local interactions adaptively. Besides, we also find that the widely used backbone in conventional CZSL settings behaves superior in perceiving local features. Thus, we construct a fusion block to combine both strengths to capture the local and non-local information. In addition, the traditional one-hot ground-truth distribution in the training phase does not reflect the accurate relationships between compositions, so we propose a composition-relation based label distribution regularization to encourage the model to actively learn the inner relationships between compositions, and extend this method to construct unseen composition pseudo distribution to further enhance the model’s generalization ability to unseen compositions. Extensive experiments and detailed analysis are conducted on three popular datasets, and the results show that our method can achieve state-of-the-art performance, which reveals its superiority in identifying novel compositions. Code is available at https://github.com/lisuyi/SFFNet_czsl.

Fusing Spatial and Frequency Features for Compositional Zero-Shot Image Classification

Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot Learning

Learning to Embed Seen/Unseen Compositions based on Graph Networks

Disentangling Before Composing: Learning Invariant Disentangled Features for Compositional Zero-Shot Learning

Dual-Stream Contrastive Learning for Compositional Zero-Shot Recognition

Learning Invariant Visual Representations for Compositional Zero-Shot Learning

Continual Compositional Zero-Shot Learning

Multi-modal Generative Adversarial Network for Zero-Shot Learning

Multidomain Features Fusion for Zero-Shot Learning

CHANNEL-WISE MIX-FUSION DEEP NEURAL NETWORKS FOR ZERO-SHOT LEARNING

Compositional Zero-Shot Learning with Contextualized Cues and Adaptive Contrastive Training

Reference-Limited Compositional Zero-Shot Learning

LVAR-CZSL: Learning Visual Attributes Representation for Compositional Zero-Shot Learning

Learning Attention as Disentangler for Compositional Zero-shot Learning

Hierarchical Prompt Learning for Compositional Zero-Shot Recognition.

Zero-Shot Compositional Concept Learning

Revealing the Proximate Long-Tail Distribution in Compositional Zero-Shot Learning

Learning object-centric complementary features for zero-shot learning

Focus-Consistent Multi-Level Aggregation for Compositional Zero-Shot Learning

Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning.

Mutual Balancing in State-Object Components for Compositional Zero-Shot Learning.