Abstract:With the rapid growth of multimedia data such as text, image, video, audio, and 3-D model, cross-media retrieval has become increasingly important, because users can retrieve the results with various types of media by submitting a query of any media type. Comparing with single-media retrieval such as image retrieval and text retrieval, cross-media retrieval is better because it provides the retrieval results with all kinds of media at the same time. In this paper, we focus on how to learn cross-media features for different media types, which is a key challenge for cross-media retrieval. Existing methods either model different media types separately or only exploit the labeled multimedia data. Actually, the data from different media types with the same semantic category are complementary to each other, and jointly modeling them is able to improve the accuracy of cross-media retrieval. In addition, although the labeled data are accurate, they require a lot of human labor and thus are very scarce. To address the above problems, we propose a semi-supervised cross-media feature learning algorithm with unified patch graph regularization (S2UPG). Our motivation and contribution mainly lie in the following three aspects. First, existing methods only model different media types in different graphs, while we employ one joint graph to simultaneously model all the media types. The joint graph is able to fully exploit the semantic correlations among various media types, which are complementary to provide the rich hint for cross-media correlation. Second, existing methods only consider the original media instances (such as images, videos, texts, audios, and 3-D models) but ignore their patches, while we make full use of both the media instances and their patches in one graph. Cross-media patches could emphasize the important parts and make cross-media correlations more precise. Third, traditional semi-supervised learning methods only exploit single-media unlabeled instances, while our approach fully exploits cross-media unlabeled instances and their patches, which can increase the diversity of training data and boost the accuracy of cross-media retrieval. Comparing with the current state-of-the-art methods on three datasets, including the challenging XMedia dataset with five media types, the comprehensive experimental results show that our proposed approach performs better.

Cross-Modality Feature Learning Via Convolutional Autoencoder

Learning Disentangled Representation for Cross-Modal Retrieval with Deep Mutual Information Estimation.

X-Gacmn: An X-Shaped Generative Adversarial Cross-Modal Network With Hypersphere Embedding

CrossMAE: Cross Modality Masked Autoencoders for Region-Aware Audio-Visual Pretraining

Learning Robust Uniform Features for Cross-Media Social Data by Using Cross Autoencoders.

Multiple Kernel Visual-Auditory Representation Learning for Retrieval

Cross-Domain Feature Learning in Multimedia

Cross-modal Semantic Autoencoder with Embedding Consensus

Poster Abstract: Representation Learning from Multimodal Sensor Data with Maximally Correlated Autoencoders

Cross-modal Image-Text Retrieval with Multitask Learning

Comprehensive Distance-Preserving Autoencoders for Cross-Modal Retrieval

Variational Autoencoder with CCA for Audio-Visual Cross-Modal Retrieval

Multimodal Masked Autoencoders Learn Transferable Representations

Unsupervised Multimodal Language Representations using Convolutional Autoencoders

Achieving Cross Modal Generalization with Multimodal Unified Representation.

Semi-Supervised Cross-Media Feature Learning with Unified Patch Graph Regularization

What Makes for Robust Multi-Modal Models in the Face of Missing Modalities?

Enhancing Separate Encoding with Multi-layer Feature Alignment for Image-Text Matching

Cross-Media Retrieval by Multimodal Representation Fusion with Deep Networks.

Dense Auto-Encoder Hashing for Robust Cross-Modality Retrieval

Learning Cross-Media Joint Representation with Sparse and Semisupervised Regularization