The Revolution of Multimodal Large Language Models: A Survey

Davide Caffagni,Federico Cocchi,Luca Barsellotti,Nicholas Moratelli,Sara Sarto,Lorenzo Baraldi,Lorenzo Baraldi,Marcella Cornia,Rita Cucchiara

2024-06-07

Abstract:Connecting text and visual modalities plays an essential role in generative intelligence. For this reason, inspired by the success of large language models, significant research efforts are being devoted to the development of Multimodal Large Language Models (MLLMs). These models can seamlessly integrate visual and textual modalities, while providing a dialogue-based interface and instruction-following capabilities. In this paper, we provide a comprehensive review of recent visual-based MLLMs, analyzing their architectural choices, multimodal alignment strategies, and training techniques. We also conduct a detailed analysis of these models across a wide range of tasks, including visual grounding, image generation and editing, visual understanding, and domain-specific applications. Additionally, we compile and describe training datasets and evaluation benchmarks, conducting comparisons among existing models in terms of performance and computational requirements. Overall, this survey offers a comprehensive overview of the current state of the art, laying the groundwork for future MLLMs.

Computer Vision and Pattern Recognition,Artificial Intelligence,Computation and Language,Multimedia

What problem does this paper attempt to address?

This paper focuses on the development of Multimodal Large Language Models (MLLMs). These models integrate visual and textual information and provide a dialogue-based interface and instruction-following capability. The paper provides a comprehensive review of recent visual-based MLLMs, analyzing their architecture choices, multimodal alignment strategies, and training techniques. The authors also analyze the performance of these models on various tasks, including visual localization, image generation and editing, visual understanding, and domain-specific applications. Additionally, they compile training datasets and evaluation benchmarks, comparing the performance and computational requirements of existing models. This survey aims to provide an overview of the current research status and lay the foundation for future developments of MLLMs. The paper particularly emphasizes the architecture of the models, training methods, and the tasks they are designed to perform.

The Revolution of Multimodal Large Language Models: A Survey

A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision-Language Tasks

Multimodal Large Language Models: A Survey

A Survey of Multimodal Large Language Model from A Data-centric Perspective

A Comprehensive Review of Multimodal Large Language Models: Performance and Challenges Across Different Tasks

A Survey on Multimodal Large Language Models

A Review of Multi-Modal Large Language and Vision Models

Efficient Multimodal Large Language Models: A Survey

Surveying the MLLM Landscape: A Meta-Review of Current Surveys

Large Multimodal Agents: A Survey

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

A Survey on Evaluation of Multimodal Large Language Models

Personalized Multimodal Large Language Models: A Survey

Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions

Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges

Visualization Literacy of Multimodal Large Language Models: A Comparative Study

LLMs Meet Multimodal Generation and Editing: A Survey

A Survey on Benchmarks of Multimodal Large Language Models

From Efficient Multimodal Models to World Models: A Survey

A Survey on Multimodal Large Language Models for Autonomous Driving

Explaining Multi-modal Large Language Models by Analyzing their Vision Perception