Abstract:The goal of cross-modal retrieval is to search for semantically similar instances in one modality by using a query from another modality. Existing approaches mainly consider the standard scenario that requires the source set for training and the target set for testing share the same scope of classes. However, they may not generalize well on zero-shot cross-modal retrieval (ZS-CMR) task, where the target set contains unseen classes that are disjoint with the seen classes in the source set. This task is more challenging due to 1) the absence of the unseen classes during training, 2) inconsistent semantics across seen and unseen classes, and 3) the heterogeneous multimodal distributions between the source and target set. To address these issues, we propose a novel Correlated Feature Synthesis and Alignment (CFSA) approach to integrate multimodal feature synthesis, common space learning and knowledge transfer for ZS-CMR. Our CFSA first utilizes class-level word embeddings to guide two coupled Wassertein generative adversarial networks (WGANs) to synthesize sufficient multimodal features with semantic correlation for stable training. Then the synthetic and true multimodal features are jointly mapped to a common semantic space via an effective distribution alignment scheme, where the cross-modal correlations of different semantic features are captured and the knowledge can be transferred to the unseen classes under the cycle-consistency constraint. Experiments on four benchmark datasets for image-text retrieval and two large-scale datasets for image-sketch retrieval show the remarkable improvements achieved by our CFAS method comparing with a bundle of state-of-the-art approaches.

Towards Improving Canonical Correlation Analysis for Cross-modal Retrieval.

3View deep canonical correlation analysis for cross-modal retrieval

Deep Canonical Correlation Analysis with Progressive and Hypergraph Learning for Cross-Modal Retrieval

A New Approach to Cross-Modal Retrieval

Two-stage deep learning for supervised cross-modal retrieval

Modeling Intra- and Inter-Pair Correlation Via Heterogeneous High-Order Preserving for Cross-Modal Retrieval

Image Retrieval Approach Based on Sparse Canonical Correlation Analysis

Cross-Modal Image Clustering Via Canonical Correlation Analysis

A Convex Discriminant Semantic Correlation Analysis for Cross-View Recognition

Incremental Canonical Correlation Analysis

Exploring Shared Subspace and Joint Sparsity for Canonical Correlation Analysis.

Intra-View and Inter-View Supervised Correlation Analysis for Multi-View Feature Learning

Canonical Principal Angles Correlation Analysis for Two-View Data

Nonnegative Constrained Graph Based Canonical Correlation Analysis for Multi-view Feature Learning

Essential of two-dimensional canonical correlation analysis and improved algorithm

A Dynamic Discriminative Canonical Correlation Analysis via Adaptive Weight Scheme

Tensor Canonical Correlation Analysis for Multi-View Dimension Reduction

Correlated Features Synthesis and Alignment for Zero-shot Cross-modal Retrieval

Cross-Modal Subspace Clustering Via Deep Canonical Correlation Analysis.

A Local Method for Canonical Correlation Analysis.

Cross‐modal Semantic Correlation Learning by Bi‐CNN Network