TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Hongyi, Cheng, Qing, Matos, Daniel, Gadi, Hari Krishna, Zhang, Yanfeng, Liu, Lu, Wang, Yongliang, Zeller, Niclas, Cremers, Daniel, Meng, Liqiu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916961450983424
author Luo, Hongyi
Cheng, Qing
Matos, Daniel
Gadi, Hari Krishna
Zhang, Yanfeng
Liu, Lu
Wang, Yongliang
Zeller, Niclas
Cremers, Daniel
Meng, Liqiu
author_facet Luo, Hongyi
Cheng, Qing
Matos, Daniel
Gadi, Hari Krishna
Zhang, Yanfeng
Liu, Lu
Wang, Yongliang
Zeller, Niclas
Cremers, Daniel
Meng, Liqiu
contents Humans can interpret geospatial information through natural language, while the geospatial cognition capabilities of Large Language Models (LLMs) remain underexplored. Prior research in this domain has been constrained by non-quantifiable metrics, limited evaluation datasets and unclear research hierarchies. Therefore, we propose a large-scale benchmark and conduct a comprehensive evaluation of the geospatial route cognition of LLMs. We create a large-scale evaluation dataset comprised of 36000 routes from 12 metropolises worldwide. Then, we introduce PathBuilder, a novel tool for converting natural language instructions into navigation routes, and vice versa, bridging the gap between geospatial information and natural language. Finally, we propose a new evaluation framework and metrics to rigorously assess 11 state-of-the-art (SOTA) LLMs on the task of route reversal. The benchmark reveals that LLMs exhibit limitation to reverse routes: most reverse routes neither return to the starting point nor are similar to the optimal route. Additionally, LLMs face challenges such as low robustness in route generation and high confidence for their incorrect answers. Code\ \&\ Data available here: \href{https://github.com/bghjmn32/EMNLP2025_Turnback}{TurnBack.}
format Preprint
id arxiv_https___arxiv_org_abs_2509_18173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route
Luo, Hongyi
Cheng, Qing
Matos, Daniel
Gadi, Hari Krishna
Zhang, Yanfeng
Liu, Lu
Wang, Yongliang
Zeller, Niclas
Cremers, Daniel
Meng, Liqiu
Machine Learning
Computation and Language
Humans can interpret geospatial information through natural language, while the geospatial cognition capabilities of Large Language Models (LLMs) remain underexplored. Prior research in this domain has been constrained by non-quantifiable metrics, limited evaluation datasets and unclear research hierarchies. Therefore, we propose a large-scale benchmark and conduct a comprehensive evaluation of the geospatial route cognition of LLMs. We create a large-scale evaluation dataset comprised of 36000 routes from 12 metropolises worldwide. Then, we introduce PathBuilder, a novel tool for converting natural language instructions into navigation routes, and vice versa, bridging the gap between geospatial information and natural language. Finally, we propose a new evaluation framework and metrics to rigorously assess 11 state-of-the-art (SOTA) LLMs on the task of route reversal. The benchmark reveals that LLMs exhibit limitation to reverse routes: most reverse routes neither return to the starting point nor are similar to the optimal route. Additionally, LLMs face challenges such as low robustness in route generation and high confidence for their incorrect answers. Code\ \&\ Data available here: \href{https://github.com/bghjmn32/EMNLP2025_Turnback}{TurnBack.}
title TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.18173