Abstract:The field of visual question answering (VQA) has seen a growing trend of integrating external knowledge sources to improve performance. However, owing to the potential incompleteness of external knowledge sources and the inherent mismatch between different forms of data, current knowledge-based visual question answering (KBVQA) techniques are still confronted with the challenge of effectively integrating and utilizing multiple heterogeneous data. To address this issue, a novel approach centered on a multi-modal semantic graph (MSG) is proposed. The MSG serves as a mechanism for effectively unifying the representation of heterogeneous data and diverse types of knowledge. Additionally, a multi-modal semantic graph knowledge reasoning model (MSG-KRM) is introduced to perform reasoning and deep fusion of image–text information and external knowledge sources. The development of the semantic graph involves extracting keywords from the image object detection information, question text, and external knowledge texts, which are then represented as symbol nodes. Three types of semantic graphs are then constructed based on the knowledge graph, including vision, question, and the external knowledge text, with non-symbol nodes added to connect these three independent graphs and marked with respective node and edge types. During the inference stage, the multi-modal semantic graph and image–text information are embedded into the feature semantic graph through three embedding methods, and a type-aware graph attention module is employed for deep reasoning. The final answer prediction is a blend of the output from the pre-trained model, graph pooling results, and the characteristics of non-symbolic nodes. The experimental results on the OK-VQA dataset show that the MSG-KRM model is superior to existing methods in terms of overall accuracy score, achieving a score of 43.58, and with improved accuracy for most subclass questions, proving the effectiveness of the proposed method.

Multimodal Knowledge Triple Extraction Based on Representation Learning

Dual Path Multi-Modal High-Order Features for Textual Content Based Visual Question Answering

Knowledge Representation Learning with Entity Descriptions, Hierarchical Types, and Textual Relations

Cross-modal Knowledge Reasoning for Knowledge-based Visual Question Answering

MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question Answering

Knowledge-Enhanced Visual Question Answering with Multi-modal Joint Guidance.

Question guided multimodal receptive field reasoning network for fact-based visual question answering

Knowledge-aware image understanding with multi-level visual representation enhancement for visual question answering

Knowledge-Based Visual Question Answering Using Multi-Modal Semantic Graph

Joint Extraction of Triple Knowledge Based on Relation Priority.

Precision Empowers, Excess Distracts: Visual Question Answering With Dynamically Infused Knowledge In Language Models

Multimodal Commonsense Knowledge Distillation for Visual Question Answering

Cross-modal Retrieval for Knowledge-based Visual Question Answering

Knowledge Condensation and Reasoning for Knowledge-based VQA

Multi-Modal Validation and Domain Interaction Learning for Knowledge-based Visual Question Answering

Visual Question Answering reasoning with external knowledge based on bimodal graph neural network

DSAMR: Dual-Stream Attention Multi-hop Reasoning for knowledge-based visual question answering

Multimodal Reranking for Knowledge-Intensive Visual Question Answering

VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question Answering

Knowledge-Embedded Mutual Guidance for Visual Reasoning

Learning Visual Knowledge Memory Networks for Visual Question Answering