M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts

Mingsheng Li,Xin Chen,Chi Zhang,Sijin Chen,Hongyuan Zhu,Fukun Yin,Gang Yu,Tao Chen

2023-12-18

Abstract:Recently, 3D understanding has become popular to facilitate autonomous agents to perform further decisionmaking. However, existing 3D datasets and methods are often limited to specific tasks. On the other hand, recent progress in Large Language Models (LLMs) and Multimodal Language Models (MLMs) have demonstrated exceptional general language and imagery tasking performance. Therefore, it is interesting to unlock MLM's potential to be 3D generalist for wider tasks. However, current MLMs' research has been less focused on 3D tasks due to a lack of large-scale 3D instruction-following datasets. In this work, we introduce a comprehensive 3D instructionfollowing dataset called M3DBench, which possesses the following characteristics: 1) It supports general multimodal instructions interleaved with text, images, 3D objects, and other visual prompts. 2) It unifies diverse 3D tasks at both region and scene levels, covering a variety of fundamental abilities in real-world 3D environments. 3) It is a large-scale 3D instruction-following dataset with over 320k instruction-response pairs. Furthermore, we establish a new benchmark for assessing the performance of large models in understanding multi-modal 3D prompts. Extensive experiments demonstrate the effectiveness of our dataset and baseline, supporting general 3D-centric tasks, which can inspire future research.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The problem this paper attempts to address is the current limitations of multimodal language models (MLMs) in handling 3D tasks. Specifically, existing 3D datasets and methods are usually limited to specific tasks, and while multimodal language models perform well on 2D vision and language tasks, research on 3D tasks is relatively scarce. The main reason for this is the lack of large-scale 3D instruction-following datasets. To overcome these limitations, the authors introduce a comprehensive 3D instruction-following dataset called M3DBench. M3DBench has the following features: 1. **Supports multimodal instructions**: The dataset supports instructions in various modalities, including text, images, and 3D objects. 2. **Unifies various 3D tasks**: It covers a range of 3D tasks from region to scene level, including visual perception, scene understanding, spatial reasoning, navigation, and planning. 3. **Large-scale dataset**: It contains over 320,000 instruction-response pairs. By constructing such a dataset, the authors aim to develop a general assistant capable of performing various tasks in real-world 3D environments and establish a new benchmark to evaluate the performance of large models in understanding multimodal 3D instructions. This will help advance future research and applications in 3D multimodal tasks.

M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts

M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions

3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset

3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding

M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

MIBench: Evaluating Multimodal Large Language Models over Multiple Images

LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark

Language-Image Models with 3D Understanding

MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

MM-BigBench: Evaluating Multimodal Models on Multimodal Content Comprehension Tasks

LLMI3D: Empowering LLM with 3D Perception from a Single 2D Image

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction

LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning

MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment

MMBench: Is Your Multi-modal Model an All-around Player?

MileBench: Benchmarking MLLMs in Long Context

3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination

Benchmarking Sequential Visual Input Reasoning and Prediction in Multimodal Large Language Models