Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Lei Li,Yuqi Wang,Runxin Xu,Peiyi Wang,Xiachong Feng,Lingpeng Kong,Qi Liu

2024-06-02

Abstract:Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes. However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a scarcity of training datasets in scientific domains. To fill this gap, we introduce Multimodal ArXiv, consisting of ArXivCap and ArXivQA, for enhancing LVLMs scientific comprehension. ArXivCap is a figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers spanning various scientific domains. Drawing from ArXivCap, we introduce ArXivQA, a question-answering dataset generated by prompting GPT-4V based on scientific figures. ArXivQA greatly enhances open-sourced LVLMs' mathematical reasoning capabilities, achieving a 10.4\% absolute accuracy gain on a multimodal mathematical reasoning benchmark. Furthermore, employing ArXivCap, we devise four vision-to-text tasks for benchmarking LVLMs. Evaluation results with state-of-the-art LVLMs underscore their struggle with the nuanced semantics of academic figures, while domain-specific training yields substantial performance gains. Our error analysis uncovers misinterpretations of visual context, recognition errors, and the production of overly simplified captions by current LVLMs, shedding light on future improvements.

Computer Vision and Pattern Recognition,Computation and Language

What problem does this paper attempt to address?

This paper focuses on the inadequate ability of large-scale vision-language models (LVLMs) to comprehend abstract charts, such as geometric shapes and scientific graphs. Due to the scarcity of training data in the scientific domain, these models perform poorly when dealing with academic chart units. To address this, the paper proposes a new dataset called Multimodal ArXiv, which consists of two subsets: ArXivCap and ArXivQA, aiming to enhance the understanding capability of LVLMs for scientific literature. ArXivCap is a scientific image-caption dataset with 6.4 million images and 3.9 million captions, collected from 572,000 interdisciplinary ArXiv papers. ArXivQA, on the other hand, contains 100,000 question-answer pairs generated based on GPT-4V prompts, which are used to improve the mathematical reasoning ability of LVLMs. Experimental results show that using ArXivQA significantly improves the accuracy of the Qwen-VL-Chat model on a multimodal mathematical reasoning benchmark test. Additionally, the paper designs four evaluation tasks based on ArXivCap to test LVLMs' understanding of academic charts. Although current LVLMs still face challenges in generating accurate titles for scientific images, training on domain-specific data can significantly improve their performance. Error analysis reveals the shortcomings of LVLMs in visual context understanding, recognition errors, and oversimplified title generation, providing directions for future improvements. In conclusion, this paper addresses how to enhance the ability of LVLMs to comprehend and process abstract charts in the scientific domain, thereby enhancing their scientific reasoning ability through the construction of new datasets and evaluation tasks.

Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding

CompCap: Improving Multimodal Large Language Models with Composite Captions

VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning

Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

Beyond the Hype: A dispassionate look at vision-language models in medical scenario

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Improving Multi-modal Large Language Model through Boosting Vision Capabilities

HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models

MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models

LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark

MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

VLIB: Unveiling insights through Visual and Linguistic Integration of Biorxiv data relevant to cancer via Multimodal Large Language Model

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI

SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation

MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding