MS2GAH: Multi-label Semantic Supervised Graph Attention Hashing for Robust Cross-Modal Retrieval

Youxiang Duan,Ning Chen,Peiying Zhang,Neeraj Kumar,Lunjie Chang,Wu Wen
DOI: https://doi.org/10.1016/j.patcog.2022.108676
IF: 8
2022-01-01
Pattern Recognition
Abstract:Due to the strong nonlinear representation capabilities of deep neural networks and the low storage and high efficiency characteristics of hash learning, deep cross-modal hashing has been propelled to the forefront of academics. How to preferably bridge semantic relevance to further bridge the semantic modality gap is the vital bottleneck to improve model performance. Confronting samples with rich semantics, how to comprehensively explore the hidden correlations and establish more precise modality relationships is the primary issue to be solved. In this work, we propose a novel deep hashing method called Multi-Label Semantic Supervised Graph Attention Hashing (MS(2)GAH), which is an end-to-end framework that integrates graph attention networks (GATs). It constructs graph features through the adjacency of nodes and assigns different weights to adjacent edges to enhance the robustness of the model. Simultaneously, multi-label annotations are utilized to bridge the semantic relevance between modalities in a more finegrained manner. To make preferable use of rich semantic information, an end-to-end label encoder is designed to mine high-level semantics from multi-label annotations to guide the feature extraction process of specific-modality networks, thereby further narrowing the modality gap. Finally, extensive experiments have been conducted on four datasets, and the results show that MS2GAH is superior to other baselines and one step forward. (C)& nbsp;2022 Elsevier Ltd. All rights reserved.
What problem does this paper attempt to address?