Abstract:The objective of visual question answering (VQA) is to adequately comprehend a question and identify relevant contents in an image that can provide an answer. Existing approaches in VQA often combine visual and question features directly to create a unified cross-modality representation for answer inference. However, this kind of approach fails to bridge the semantic gap between visual and text modalities, resulting in a lack of alignment in cross-modality semantics and the inability to match key visual content accurately. In this article, we propose a model called the caption bridge-based cross-modality alignment and contrastive learning model (CBAC) to address the issue. The CBAC model aims to reduce the semantic gap between different modalities. It consists of a caption-based cross-modality alignment module and a visual-caption (V-C) contrastive learning module. By utilizing an auxiliary caption that shares the same modality as the question and has closer semantic associations with the visual, we are able to effectively reduce the semantic gap by separately matching the caption with both the question and the visual to generate pre-alignment features for each, which are then used in the subsequent fusion process. We also leverage the fact that V-C pairs exhibit stronger semantic connections compared to question-visual (Q-V) pairs to employ a contrastive learning mechanism on visual and caption pairs to further enhance the semantic alignment capabilities of single-modality encoders. Extensive experiments conducted on three benchmark datasets demonstrate that the proposed model outperforms previous state-of-the-art VQA models. Additionally, ablation experiments confirm the effectiveness of each module in our model. Furthermore, we conduct a qualitative analysis by visualizing the attention matrices to assess the reasoning reliability of the proposed model.

CausalME: Balancing bi-modalities in Visual Question Answering.

Dual Path Multi-Modal High-Order Features for Textual Content Based Visual Question Answering

Simple and Effective Visual Question Answering in a Single Modality

Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective

Reducing Vision-Answer Biases for Multiple-Choice VQA

Debiasing Medical Visual Question Answering via Counterfactual Training

Collaborative Modality Fusion for Mitigating Language Bias in Visual Question Answering

Unveiling Cross Modality Bias in Visual Question Answering: A Causal View with Possible Worlds VQA

Overcoming Language Priors In Vqa Via Decomposed Linguistic Representations

Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention

Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models

Removing Bias of Video Question Answering by Causal Theory

An Empirical Study on the Language Modal in Visual Question Answering

From Superficial to Deep: Language Bias Driven Curriculum Learning for Visual Question Answering.

Visual Causal Scene Refinement for Video Question Answering

Towards Robust Visual Question Answering: Making the Most of Biased Samples via Contrastive Learning

Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering

Bridging the Cross-Modality Semantic Gap in Visual Question Answering

SC-ML: Self-supervised Counterfactual Metric Learning for Debiased Visual Question Answering

Causal Reasoning through Two Layers of Cognition for Improving Generalization in Visual Question Answering

Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality