Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

Chuofan Ma,Yi Jiang,Jiannan Wu,Zehuan Yuan,Xiaojuan Qi

2024-04-20

Abstract:We introduce Groma, a Multimodal Large Language Model (MLLM) with grounded and fine-grained visual perception ability. Beyond holistic image understanding, Groma is adept at region-level tasks such as region captioning and visual grounding. Such capabilities are built upon a localized visual tokenization mechanism, where an image input is decomposed into regions of interest and subsequently encoded into region tokens. By integrating region tokens into user instructions and model responses, we seamlessly enable Groma to understand user-specified region inputs and ground its textual output to images. Besides, to enhance the grounded chat ability of Groma, we curate a visually grounded instruction dataset by leveraging the powerful GPT-4V and visual prompting techniques. Compared with MLLMs that rely on the language model or external module for localization, Groma consistently demonstrates superior performances in standard referring and grounding benchmarks, highlighting the advantages of embedding localization into image tokenization. Project page:

Computer Vision and Pattern Recognition,Artificial Intelligence,Computation and Language,Machine Learning

What problem does this paper attempt to address?

The paper proposes Groma, a multimodal large-scale language model with fine-grained visual understanding capability. Groma excels in region-level tasks such as region description and visual localization. To address the lack of localization ability in existing models, Groma introduces a local visual token mechanism that decomposes images into regions of interest and encodes them as region tokens. By integrating the region tokens into user instructions and model responses, Groma is able to understand region-specific inputs from users and associate its text output with the image. Furthermore, to enhance Groma's conversational ability, they create a visual localization instruction dataset using GPT-4V and visual prompting techniques. Compared to methods relying on language models or external modules for localization, Groma performs impressively on standard citation and localization benchmarks, demonstrating the advantages of embedding localization into image tokens.

Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

GROUNDHOG: Grounding Large Language Models to Holistic Segmentation

GLaMM: Pixel Grounding Large Multimodal Model

GroundingGPT:Language Enhanced Multi-modal Grounding Model

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Pixel Aligned Language Models

LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models

LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding

InfMLLM: A Unified Framework for Visual-Language Tasks.

LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent

Grounded 3D-LLM with Referent Tokens

What Makes for Good Visual Tokenizers for Large Language Models?

Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models

Efficient Multi-modal Large Language Models via Visual Token Grouping

Emerging Pixel Grounding in Large Multimodal Models Without Grounding Supervision

Learning Visual Grounding from Generative Vision and Language Model

Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Transformer-based Visual Grounding with Cross-modality Interaction

Visual Grounding With Joint Multimodal Representation and Interaction