MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dihan, Mahir Labib, Hassan, Md Tanvir, Parvez, Md Tanvir, Hasan, Md Hasebul, Alam, Md Almash, Cheema, Muhammad Aamir, Ali, Mohammed Eunus, Parvez, Md Rizwan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912416747487232
author Dihan, Mahir Labib
Hassan, Md Tanvir
Parvez, Md Tanvir
Hasan, Md Hasebul
Alam, Md Almash
Cheema, Muhammad Aamir
Ali, Mohammed Eunus
Parvez, Md Rizwan
author_facet Dihan, Mahir Labib
Hassan, Md Tanvir
Parvez, Md Tanvir
Hasan, Md Hasebul
Alam, Md Almash
Cheema, Muhammad Aamir
Ali, Mohammed Eunus
Parvez, Md Rizwan
contents Recent advancements in foundation models have improved autonomous tool usage and reasoning, but their capabilities in map-based reasoning remain underexplored. To address this, we introduce MapEval, a benchmark designed to assess foundation models across three distinct tasks - textual, API-based, and visual reasoning - through 700 multiple-choice questions spanning 180 cities and 54 countries, covering spatial relationships, navigation, travel planning, and real-world map interactions. Unlike prior benchmarks that focus on simple location queries, MapEval requires models to handle long-context reasoning, API interactions, and visual map analysis, making it the most comprehensive evaluation framework for geospatial AI. On evaluation of 30 foundation models, including Claude-3.5-Sonnet, GPT-4o, and Gemini-1.5-Pro, none surpass 67% accuracy, with open-source models performing significantly worse and all models lagging over 20% behind human performance. These results expose critical gaps in spatial inference, as models struggle with distances, directions, route planning, and place-specific reasoning, highlighting the need for better geospatial AI to bridge the gap between foundation models and real-world navigation. All the resources are available at: https://mapeval.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00316
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models
Dihan, Mahir Labib
Hassan, Md Tanvir
Parvez, Md Tanvir
Hasan, Md Hasebul
Alam, Md Almash
Cheema, Muhammad Aamir
Ali, Mohammed Eunus
Parvez, Md Rizwan
Computation and Language
Recent advancements in foundation models have improved autonomous tool usage and reasoning, but their capabilities in map-based reasoning remain underexplored. To address this, we introduce MapEval, a benchmark designed to assess foundation models across three distinct tasks - textual, API-based, and visual reasoning - through 700 multiple-choice questions spanning 180 cities and 54 countries, covering spatial relationships, navigation, travel planning, and real-world map interactions. Unlike prior benchmarks that focus on simple location queries, MapEval requires models to handle long-context reasoning, API interactions, and visual map analysis, making it the most comprehensive evaluation framework for geospatial AI. On evaluation of 30 foundation models, including Claude-3.5-Sonnet, GPT-4o, and Gemini-1.5-Pro, none surpass 67% accuracy, with open-source models performing significantly worse and all models lagging over 20% behind human performance. These results expose critical gaps in spatial inference, as models struggle with distances, directions, route planning, and place-specific reasoning, highlighting the need for better geospatial AI to bridge the gap between foundation models and real-world navigation. All the resources are available at: https://mapeval.github.io/.
title MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models
topic Computation and Language
url https://arxiv.org/abs/2501.00316