SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Jiabin, Wang, Haiping, Li, Jinpeng, Liu, Yuan, Dong, Zhen, Yang, Bisheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909624020500480
author Chen, Jiabin
Wang, Haiping
Li, Jinpeng
Liu, Yuan
Dong, Zhen
Yang, Bisheng
author_facet Chen, Jiabin
Wang, Haiping
Li, Jinpeng
Liu, Yuan
Dong, Zhen
Yang, Bisheng
contents We propose SpatialLLM, a novel approach advancing spatial intelligence tasks in complex urban scenes. Unlike previous methods requiring geographic analysis tools or domain expertise, SpatialLLM is a unified language model directly addressing various spatial intelligence tasks without any training, fine-tuning, or expert intervention. The core of SpatialLLM lies in constructing detailed and structured scene descriptions from raw spatial data to prompt pre-trained LLMs for scene-based analysis. Extensive experiments show that, with our designs, pretrained LLMs can accurately perceive spatial distribution information and enable zero-shot execution of advanced spatial intelligence tasks, including urban planning, ecological analysis, traffic management, etc. We argue that multi-field knowledge, context length, and reasoning ability are key factors influencing LLM performances in urban analysis. We hope that SpatialLLM will provide a novel viable perspective for urban intelligent analysis and management. The code and dataset are available at https://github.com/WHU-USI3DV/SpatialLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12703
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
Chen, Jiabin
Wang, Haiping
Li, Jinpeng
Liu, Yuan
Dong, Zhen
Yang, Bisheng
Computer Vision and Pattern Recognition
Artificial Intelligence
We propose SpatialLLM, a novel approach advancing spatial intelligence tasks in complex urban scenes. Unlike previous methods requiring geographic analysis tools or domain expertise, SpatialLLM is a unified language model directly addressing various spatial intelligence tasks without any training, fine-tuning, or expert intervention. The core of SpatialLLM lies in constructing detailed and structured scene descriptions from raw spatial data to prompt pre-trained LLMs for scene-based analysis. Extensive experiments show that, with our designs, pretrained LLMs can accurately perceive spatial distribution information and enable zero-shot execution of advanced spatial intelligence tasks, including urban planning, ecological analysis, traffic management, etc. We argue that multi-field knowledge, context length, and reasoning ability are key factors influencing LLM performances in urban analysis. We hope that SpatialLLM will provide a novel viable perspective for urban intelligent analysis and management. The code and dataset are available at https://github.com/WHU-USI3DV/SpatialLLM.
title SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.12703