Abstract:Generative foundation models have advanced large-scale text-driven natural image generation, becoming a prominent research trend across various vertical domains. However, in the remote sensing field, there is still a lack of research on large-scale text-to-image (text2image) generation technology. Existing remote sensing image-text datasets are small in scale and confined to specific geographic areas and scene types. Besides, existing text2image methods have struggled to achieve global-scale, multi-resolution controllable, and unbounded image generation. To address these challenges, this paper presents two key contributions: the Git-10M dataset and the Text2Earth foundation model. Git-10M is a global-scale image-text dataset comprising 10 million image-text pairs, 5 times larger than the previous largest one. The dataset covers a wide range of geographic scenes and contains resolution information, significantly surpassing existing datasets in both size and diversity. Building on Git-10M, we propose Text2Earth, a 1.3 billion parameter generative foundation model based on the diffusion framework to model global-scale remote sensing scenes. Text2Earth integrates a resolution guidance mechanism, enabling users to specify image resolutions. A dynamic condition adaptation strategy is proposed for training and inference to improve image quality. Text2Earth excels in zero-shot text2image generation and demonstrates robust generalization and flexibility across multiple tasks, including unbounded scene construction, image editing, and cross-modal image generation. This robust capability surpasses previous models restricted to the basic fixed size and limited scene types. On the previous benchmark dataset, Text2Earth outperforms previous models with an improvement of +26.23 FID and +20.95% Zero-shot Cls-OA <a class="link-external link-http" href="http://metric.Our" rel="external noopener nofollow">this http URL</a> project page is \url{<a class="link-external link-https" href="https://chen-yang-liu.github.io/Text2Earth" rel="external noopener nofollow">this https URL</a>}

Text-space Graph Foundation Models: Comprehensive Benchmarks and New Insights

Position: Graph Foundation Models are Already Here

GraphFM: A Comprehensive Benchmark for Graph Foundation Model

GFT: Graph Foundation Model with Transferable Tree Vocabulary

LangGFM: A Large Language Model Alone Can be a Powerful Graph Foundation Model

Towards Graph Foundation Models: A Survey and Beyond

DTGB: A Comprehensive Benchmark for Dynamic Text-Attributed Graphs

Lecture-style Tutorial: Towards Graph Foundation Models

TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models

OpenFGL: A Comprehensive Benchmarks for Federated Graph Learning

GOFA: A Generative One-For-All Model for Joint Graph Language Modeling

GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models

TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs

PANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation Models

OpenGSL: A Comprehensive Benchmark for Graph Structure Learning

Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model

Multimodal Graph Benchmark

Temporal Graph Benchmark for Machine Learning on Temporal Graphs

Text-Free Multi-domain Graph Pre-training: Toward Graph Foundation Models

T$^3$Bench: Benchmarking Current Progress in Text-to-3D Generation

UGSL: A Unified Framework for Benchmarking Graph Structure Learning