Abstract:Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly represent advancements, thereby impeding further progress in the field. Current evaluations heavily rely on classification and caption tasks, falling short in providing a thorough assessment of MLLMs. A pressing need exists for a more sophisticated evaluation method capable of thoroughly analyzing the spatial understanding and expressive capabilities of these models. To address these issues, we introduce a scalable 3D benchmark, accompanied by a large-scale instruction-tuning dataset known as 3DBench, providing an extensible platform for a comprehensive evaluation of MLLMs. Specifically, we establish the benchmark that spans a wide range of spatial and semantic scales, from object-level to scene-level, addressing both perception and planning tasks. Furthermore, we present a rigorous pipeline for automatically constructing scalable 3D instruction-tuning datasets, covering 10 diverse multi-modal tasks with more than 0.23 million QA pairs generated in total. Thorough experiments evaluating trending MLLMs, comparisons against existing datasets, and variations of training protocols demonstrate the superiority of 3DBench, offering valuable insights into current limitations and potential research directions.

What problem does this paper attempt to address?

This paper focuses on the challenges in evaluating multimodal large-scale language models (MLLMs) that combine point clouds and language. The current evaluation methods mainly rely on classification and captioning tasks, but fail to comprehensively assess the spatial understanding and expression capabilities of MLLMs. To address this issue, the paper proposes an scalable 3D benchmark test platform called 3DBench, as well as a large-scale instruction fine-tuning dataset for comprehensive evaluation of these models. 3DBench benchmark covers various spatial and semantic scales ranging from object-level to scene-level, including perception and planning tasks. In addition, the paper introduces an automated method to construct a large-scale 3D instruction fine-tuning dataset, which covers 10 different multimodal tasks with over 230,000 QA pairs in total. Experimental results demonstrate that 3DBench outperforms existing datasets, providing insights into current limitations and future research directions. Existing 3D instruction fine-tuning datasets mainly come from public datasets, but lack effective evaluation of models' spatial understanding capabilities for complex and scene-oriented point clouds. 3DBench addresses this problem by expanding the task types and introducing customized evaluation metrics. It includes metrics for text generation quality evaluation, traditional accuracy and IOU metrics, GPT scores, and a new path loss metric. Furthermore, the 3DBench dataset is generated through an automated process that extracts metadata and reconstructs 3D objects and scenes from the Procthor simulation framework, resulting in over 2.3 million instruction fine-tuning samples. Through experiments, the authors validate the effectiveness of 3DBench and compare it with other models and datasets to highlight its advantages in evaluating 3D-LLMs. In summary, this paper aims to establish a more comprehensive evaluation standard to promote further development in the 3D domain by providing a platform that enables in-depth analysis of the model's spatial understanding and expression capabilities.

3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset

M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts

M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions

3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding

LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark

MIBench: Evaluating Multimodal Large Language Models over Multiple Images

MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark

MMBench: Is Your Multi-modal Model an All-around Player?

MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models

MileBench: Benchmarking MLLMs in Long Context

Q-Bench+: A Benchmark for Multi-modal Foundation Models on Low-level Vision from Single Images to Pairs

UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios

Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study

MM-BigBench: Evaluating Multimodal Models on Multimodal Content Comprehension Tasks

M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

A Survey on Multimodal Benchmarks: In the Era of Large AI Models

A Survey on Benchmarks of Multimodal Large Language Models

T$^3$Bench: Benchmarking Current Progress in Text-to-3D Generation

MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs

3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI