Coupled Feature Mapping and Correlation Mining for Cross-Media Retrieval

Mengdi Fan,Wenmin Wang,Ronggang Wang
DOI: https://doi.org/10.1109/icmew.2016.7574754
2016-01-01
Abstract:Cross-media retrieval aims to integrate and analyze the features of various modalities (e.g., text, image and video) to mine their potential semantic information. In this paper, we propose a novel cross-media retrieval framework, which performs coupled feature mapping and correlation mining successively. Our method first learns two projection matrices to map the multimodal features into a common category space, in which homo- and hetero-correlation techniques can be applied easily. Homo-correlation focuses on the semantic category information within the same media type, while hetero-correlation focuses on the semantic category information be-tween different media types. The two could complement and reinforce each other. Experiments on two different datasets, Wikipedia dataset and Pascal Voc dataset, demonstrate that the proposed framework gives promising results compared to the related state-of-the-art approaches.
What problem does this paper attempt to address?