Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roberts, Jonathan, Lüddecke, Timo, Sheikh, Rehan, Han, Kai, Albanie, Samuel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914642518867968
author Roberts, Jonathan
Lüddecke, Timo
Sheikh, Rehan
Han, Kai
Albanie, Samuel
author_facet Roberts, Jonathan
Lüddecke, Timo
Sheikh, Rehan
Han, Kai
Albanie, Samuel
contents Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains are yet to be explored, despite potential wide-ranging benefits to navigation, environmental research, urban development, and disaster response. We conduct a series of experiments exploring various vision capabilities of MLLMs within these domains, particularly focusing on the frontier model GPT-4V, and benchmark its performance against open-source counterparts. Our methodology involves challenging these models with a small-scale geographic benchmark consisting of a suite of visual tasks, testing their abilities across a spectrum of complexity. The analysis uncovers not only where such models excel, including instances where they outperform humans, but also where they falter, providing a balanced view of their capabilities in the geographic domain. To enable the comparison and evaluation of future models, our benchmark will be publicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2311_14656
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
Roberts, Jonathan
Lüddecke, Timo
Sheikh, Rehan
Han, Kai
Albanie, Samuel
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains are yet to be explored, despite potential wide-ranging benefits to navigation, environmental research, urban development, and disaster response. We conduct a series of experiments exploring various vision capabilities of MLLMs within these domains, particularly focusing on the frontier model GPT-4V, and benchmark its performance against open-source counterparts. Our methodology involves challenging these models with a small-scale geographic benchmark consisting of a suite of visual tasks, testing their abilities across a spectrum of complexity. The analysis uncovers not only where such models excel, including instances where they outperform humans, but also where they falter, providing a balanced view of their capabilities in the geographic domain. To enable the comparison and evaluation of future models, our benchmark will be publicly released.
title Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2311.14656