Abstract:Medical Visual Question Answering (VQA) is an important task in medical multi-modal Large Language Models (LLMs), aiming to answer clinically relevant questions regarding input medical images. This technique has the potential to improve the efficiency of medical professionals while relieving the burden on the public health system, particularly in resource-poor countries. However, existing medical VQA datasets are small and only contain simple questions (equivalent to classification tasks), which lack semantic reasoning and clinical knowledge. Our previous work proposed a clinical knowledge-driven image difference VQA benchmark using a rule-based approach (Hu et al., 2023). However, given the same breadth of information coverage, the rule-based approach shows an 85% error rate on extracted labels. We trained an LLM method to extract labels with 62% increased accuracy. We also comprehensively evaluated our labels with 2 clinical experts on 100 samples to help us fine-tune the LLM. Based on the trained LLM model, we proposed a large-scale medical VQA dataset, Medical-CXR-VQA, using LLMs focused on chest X-ray images. The questions involved detailed information, such as abnormalities, locations, levels, and types. Based on this dataset, we proposed a novel VQA method by constructing three different relationship graphs: spatial relationships, semantic relationships, and implicit relationship graphs on the image regions, questions, and semantic labels. We leveraged graph attention to learn the logical reasoning paths for different questions. These learned graph VQA reasoning paths can be further used for LLM prompt engineering and chain-of-thought, which are crucial for further fine-tuning and training multi-modal large language models. Moreover, we demonstrate that our approach has the qualities of evidence and faithfulness, which are crucial in the clinical field. The code and the dataset is available at https://github.com/Holipori/Medical-CXR-VQA.

SYSU-HCP at VQA-Med 2021: A Data-centric Model with Efficient Training Methodology for Medical Visual Question Answering.

Dual Path Multi-Modal High-Order Features for Textual Content Based Visual Question Answering

HCP-MIC at VQA-Med 2020: Effective Visual Representation for Medical Visual Question Answering.

Question-guided Feature Pyramid Network for Medical Visual Question Answering

TAM at VQA-Med 2021 - A Hybrid Model with Feature Extraction and Fusion for Medical Visual Question Answering.

CGMVQA: A New Classification and Generative Model for Medical Visual Question Answering

Hierarchical deep multi-modal network for medical visual question answering

Medical visual question answering via corresponding feature fusion combined with semantic attention

MHKD-MVQA: Multimodal Hierarchical Knowledge Distillation for Medical Visual Question Answering.

Medical visual question answering with symmetric interaction attention and cross-modal gating

PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

MF2-MVQA: A Multi-stage Feature Fusion method for Medical Visual Question Answering

Masked Vision and Language Pre-training with Unimodal and Multimodal Contrastive Losses for Medical Visual Question Answering

Medical visual question answering using joint self-supervised learning

Asymmetric cross-modal attention network with multimodal augmented mixup for medical visual question answering

A Bi-level representation learning model for medical visual question answering

Towards Visual Question Answering on Pathology Images.

Medical visual question answering based on question-type reasoning and semantic space constraint

Cross-Modal Self-Supervised Vision Language Pre-training with Multiple Objectives for Medical Visual Question Answering

Interpretable medical image Visual Question Answering via multi-modal relationship graph learning

Caption-Aware Medical VQA Via Semantic Focusing and Progressive Cross-Modality Comprehension