Abstract:Recently, large language models (LLMs) with extensive general knowledge and powerful reasoning abilities have seen rapid development and widespread application. A systematic and reliable evaluation of LLMs or vision-language model (VLMs) is a crucial step in applying and developing them for various fields. There have been some early explorations about the usability of LLMs for limited urban tasks, but a systematic and scalable evaluation benchmark is still lacking. The challenge in constructing a systematic evaluation benchmark for urban research lies in the diversity of urban data, the complexity of application scenarios and the highly dynamic nature of the urban environment. In this paper, we design CityBench, an interactive simulator based evaluation platform, as the first systematic benchmark for evaluating the capabilities of LLMs for diverse tasks in urban research. First, we build CityData to integrate the diverse urban data and CitySimu to simulate fine-grained urban dynamics. Based on CityData and CitySimu, we design 8 representative urban tasks in 2 categories of perception-understanding and decision-making as the CityBench. With extensive results from 30 well-known LLMs and VLMs in 13 cities around the world, we find that advanced LLMs and VLMs can achieve competitive performance in diverse urban tasks requiring commonsense and semantic understanding abilities, e.g., understanding the human dynamics and semantic inference of urban images. Meanwhile, they fail to solve the challenging urban tasks requiring professional knowledge and high-level reasoning abilities, e.g., geospatial prediction and traffic control task. These observations provide valuable perspectives for utilizing and developing LLMs in the future. Codes are openly accessible via <a class="link-external link-https" href="https://github.com/tsinghua-fib-lab/CityBench" rel="external noopener nofollow">this https URL</a>.

SysBench: Can Large Language Models Follow System Messages?

CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery

FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

CFBench: A Comprehensive Constraints-Following Benchmark for LLMs

Benchmarking the Text-to-SQL Capability of Large Language Models: A Comprehensive Evaluation

Towards Benchmarking Situational Awareness of Large Language Models:Comprehensive Benchmark, Evaluation and Analysis

Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

AgentBench: Evaluating LLMs as Agents

A Survey on Benchmarks of Multimodal Large Language Models

ML-Bench: Large Language Models Leverage Open-source Libraries for Machine Learning Tasks

CPSDBench: A Large Language Model Evaluation Benchmark and Baseline for Chinese Public Security Domain

CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models

CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks

LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context Scenarios

FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

NLPBench: Evaluating Large Language Models on Solving NLP Problems

Evaluating Large Language Models at Evaluating Instruction Following

MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

ZhuJiu: A Multi-dimensional, Multi-faceted Chinese Benchmark for Large Language Models

SimulBench: Evaluating Language Models with Creative Simulation Tasks

Benchmarking Foundation Models with Language-Model-as-an-Examiner