Sat2Cap: Mapping Fine-Grained Textual Descriptions from Satellite Images

Aayush Dhakal,Adeel Ahmad,Subash Khanal,Srikumar Sastry,Hannah Kerner,Nathan Jacobs

2024-04-12

Abstract:We propose a weakly supervised approach for creating maps using free-form textual descriptions. We refer to this work of creating textual maps as zero-shot mapping. Prior works have approached mapping tasks by developing models that predict a fixed set of attributes using overhead imagery. However, these models are very restrictive as they can only solve highly specific tasks for which they were trained. Mapping text, on the other hand, allows us to solve a large variety of mapping problems with minimal restrictions. To achieve this, we train a contrastive learning framework called Sat2Cap on a new large-scale dataset with 6.1M pairs of overhead and ground-level images. For a given location and overhead image, our model predicts the expected CLIP embeddings of the ground-level scenery. The predicted CLIP embeddings are then used to learn about the textual space associated with that location. Sat2Cap is also conditioned on date-time information, allowing it to model temporally varying concepts over a location. Our experimental results demonstrate that our models successfully capture ground-level concepts and allow large-scale mapping of fine-grained textual queries. Our approach does not require any text-labeled data, making the training easily scalable. The code, dataset, and models will be made publicly available.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The problem this paper attempts to address is how to generate fine-grained textual descriptions from satellite images, specifically by creating a method that can map satellite images to detailed text descriptions. Traditional methods typically rely on training models for specific tasks to predict a fixed set of attributes, which limits their applicability. This paper proposes a weakly supervised approach that learns a semantically rich embedding space between geographic locations and fine-grained text, enabling zero-shot map generation. This method allows users to generate maps through natural language queries without the need to train models for each specific task, thus providing greater flexibility and broader application possibilities. Specifically, the main contributions of the paper include: 1. **Weakly Supervised Method**: A weakly supervised method is proposed for learning fine-grained textual concepts of geographic locations. 2. **Zero-Shot Map Generation**: A method for generating large-scale maps from text queries is achieved, as shown in Figure 1. 3. **New Large-Scale Cross-View Dataset**: A new dataset containing 6.1 million pairs of top-view and ground images is created. Through these contributions, the paper aims to overcome the limitations of existing methods and provide a more general and flexible map generation framework.

Sat2Cap: Mapping Fine-Grained Textual Descriptions from Satellite Images

SSC: Semantic Scan Context for Large-Scale Place Recognition

From Satellite to Ground: Satellite Assisted Visual Localization with Cross-view Semantic Matching

ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models

Discoverability in Satellite Imagery: A Good Sentence is Worth a Thousand Pictures

SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

Generating Spatial-aware Captions for TextCaps

TextCaps: a Dataset for Image Captioning with Reading Comprehension

SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting

Visual Spatial Description: Controlled Spatial-Oriented Image-to-Text Generation

Attentive Weakly Supervised land cover mapping for object-based satellite image time series data with spatial interpretation

SATIN: A Multi-Task Metadataset for Classifying Satellite Imagery using Vision-Language Models

SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing

OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map Construction

Scalable 3D Captioning with Pretrained Models

Tag Map: A Text-Based Map for Spatial Reasoning and Navigation with Large Language Models

Deep Learning for Understanding Satellite Imagery: An Experimental Survey

An Efficient System for Automatic Map Storytelling -- A Case Study on Historical Maps

HDMapNet: A Local Semantic Map Learning and Evaluation Framework.

Semantic Annotation of High-Resolution Satellite Images Via Weakly Supervised Learning.

Learning Tri-modal Embeddings for Zero-Shot Soundscape Mapping