Abstract:Cross-modality human behavior analysis has attracted much attention from both academia and industry. In this article, we focus on the cross-modality image-text retrieval problem for human behavior analysis, which can learn a common latent space for cross-modality data and thus benefit the understanding of human behavior with data from different modalities. Existing state-of-the-art cross-modality image-text retrieval models tend to be fine-grained region-word matching approaches, where they begin with measuring similarities for each image region or text word followed by aggregating them to estimate the global image-text similarity. However, it is observed that such fine-grained approaches often encounter the similarity bias problem, because they only consider matched text words for an image region or matched image regions for a text word for similarity calculation, but they totally ignore unmatched words/regions, which might still be salient enough to affect the global image-text similarity. In this article, we propose an Adaptive Confidence Matching Network (ACMNet), which is also a fine-grained matching approach, to effectively deal with such a similarity bias. Apart from calculating the local similarity for each region(/word) with its matched words(/regions), ACMNet also introduces a confidence score for the local similarity by leveraging the global text(/image) information, which is expected to help measure the semantic relatedness of the region(/word) to the whole text(/image). Moreover, ACMNet also incorporates the confidence scores together with the local similarities in estimating the global image-text similarity. To verify the effectiveness of ACMNet, we conduct extensive experiments and make comparisons with state-of-the-art methods on two benchmark datasets, i.e., Flickr30k and MS COCO. Experimental results show that the proposed ACMNet can outperform the state-of-the-art methods by a clear margin, which well demonstrates the effectiveness of the proposed ACMNet in human behavior analysis and the reasonableness of tackling the mentioned similarity bias issue.

Cross-Modal Knowledge Adaptation for Language-Based Person Search

ACMNet

Text-based person search via cross-modal alignment learning

Hierarchical Gumbel Attention Network for Text-based Person Search

A Cross-modality and Progressive Person Search System

Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval

Hybrid Attention Network for Language-Based Person Search

SCMM: Calibrating Cross-modal Representations for Text-Based Person Search

Top-Push Constrained Modality-Adaptive Dictionary Learning for Cross-Modality Person Re-Identification

Cross-Modal Knowledge Discovery, Inference, and Challenges.

CUHK at ImageCLEF 2005: cross-language and cross-media image retrieval

CMPD: Using Cross Memory Network With Pair Discrimination for Image-Text Retrieval

Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

Pose-Guided Multi-Granularity Attention Network for Text-Based Person Search

Adaptive Uncertainty-Based Learning for Text-Based Person Retrieval

TIPCB: A simple but effective part-based convolutional baseline for text-based person search

Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search Benchmark

Unpaired Image-text Matching via Multimodal Aligned Conceptual Knowledge

CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval

Cross‐modal knowledge learning with scene text for fine‐grained image classification

Adaptive Cross-Modal Prototypes for Cross-Domain Visual-Language Retrieval