StreetviewLLM: Extracting Geographic Information Using a Chain-of-Thought Multimodal Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zongrong, Xu, Junhao, Wang, Siqin, Wu, Yifan, Li, Haiyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915030684925952
author Li, Zongrong
Xu, Junhao
Wang, Siqin
Wu, Yifan
Li, Haiyang
author_facet Li, Zongrong
Xu, Junhao
Wang, Siqin
Wu, Yifan
Li, Haiyang
contents Geospatial predictions are crucial for diverse fields such as disaster management, urban planning, and public health. Traditional machine learning methods often face limitations when handling unstructured or multi-modal data like street view imagery. To address these challenges, we propose StreetViewLLM, a novel framework that integrates a large language model with the chain-of-thought reasoning and multimodal data sources. By combining street view imagery with geographic coordinates and textual data, StreetViewLLM improves the precision and granularity of geospatial predictions. Using retrieval-augmented generation techniques, our approach enhances geographic information extraction, enabling a detailed analysis of urban environments. The model has been applied to seven global cities, including Hong Kong, Tokyo, Singapore, Los Angeles, New York, London, and Paris, demonstrating superior performance in predicting urban indicators, including population density, accessibility to healthcare, normalized difference vegetation index, building height, and impervious surface. The results show that StreetViewLLM consistently outperforms baseline models, offering improved predictive accuracy and deeper insights into the built environment. This research opens new opportunities for integrating the large language model into urban analytics, decision-making in urban planning, infrastructure management, and environmental monitoring.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14476
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StreetviewLLM: Extracting Geographic Information Using a Chain-of-Thought Multimodal Large Language Model
Li, Zongrong
Xu, Junhao
Wang, Siqin
Wu, Yifan
Li, Haiyang
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Geospatial predictions are crucial for diverse fields such as disaster management, urban planning, and public health. Traditional machine learning methods often face limitations when handling unstructured or multi-modal data like street view imagery. To address these challenges, we propose StreetViewLLM, a novel framework that integrates a large language model with the chain-of-thought reasoning and multimodal data sources. By combining street view imagery with geographic coordinates and textual data, StreetViewLLM improves the precision and granularity of geospatial predictions. Using retrieval-augmented generation techniques, our approach enhances geographic information extraction, enabling a detailed analysis of urban environments. The model has been applied to seven global cities, including Hong Kong, Tokyo, Singapore, Los Angeles, New York, London, and Paris, demonstrating superior performance in predicting urban indicators, including population density, accessibility to healthcare, normalized difference vegetation index, building height, and impervious surface. The results show that StreetViewLLM consistently outperforms baseline models, offering improved predictive accuracy and deeper insights into the built environment. This research opens new opportunities for integrating the large language model into urban analytics, decision-making in urban planning, infrastructure management, and environmental monitoring.
title StreetviewLLM: Extracting Geographic Information Using a Chain-of-Thought Multimodal Large Language Model
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.14476