Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Liuchang, Zhao, Shuo, Lin, Qingming, Chen, Luyao, Luo, Qianqian, Wu, Sensen, Ye, Xinyue, Feng, Hailin, Du, Zhenhong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915088670130176
author Xu, Liuchang
Zhao, Shuo
Lin, Qingming
Chen, Luyao
Luo, Qianqian
Wu, Sensen
Ye, Xinyue
Feng, Hailin
Du, Zhenhong
author_facet Xu, Liuchang
Zhao, Shuo
Lin, Qingming
Chen, Luyao
Luo, Qianqian
Wu, Sensen
Ye, Xinyue
Feng, Hailin
Du, Zhenhong
contents The emergence of large language models such as ChatGPT, Gemini, and others highlights the importance of evaluating their diverse capabilities, ranging from natural language understanding to code generation. However, their performance on spatial tasks has not been thoroughly assessed. This study addresses this gap by introducing a new multi-task spatial evaluation dataset designed to systematically explore and compare the performance of several advanced models on spatial tasks. The dataset includes twelve distinct task types, such as spatial understanding and simple route planning, each with verified and accurate answers. We evaluated multiple models, including OpenAI's gpt-3.5-turbo, gpt-4-turbo, gpt-4o, ZhipuAI's glm-4, Anthropic's claude-3-sonnet-20240229, and MoonShot's moonshot-v1-8k, using a two-phase testing approach. First, we conducted zero-shot testing. Then, we categorized the dataset by difficulty and performed prompt-tuning tests. Results show that gpt-4o achieved the highest overall accuracy in the first phase, with an average of 71.3%. Although moonshot-v1-8k slightly underperformed overall, it outperformed gpt-4o in place name recognition tasks. The study also highlights the impact of prompt strategies on model performance in specific tasks. For instance, the Chain-of-Thought (CoT) strategy increased gpt-4o's accuracy in simple route planning from 12.4% to 87.5%, while a one-shot strategy improved moonshot-v1-8k's accuracy in mapping tasks from 10.1% to 76.3%.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14438
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study
Xu, Liuchang
Zhao, Shuo
Lin, Qingming
Chen, Luyao
Luo, Qianqian
Wu, Sensen
Ye, Xinyue
Feng, Hailin
Du, Zhenhong
Computation and Language
Computers and Society
The emergence of large language models such as ChatGPT, Gemini, and others highlights the importance of evaluating their diverse capabilities, ranging from natural language understanding to code generation. However, their performance on spatial tasks has not been thoroughly assessed. This study addresses this gap by introducing a new multi-task spatial evaluation dataset designed to systematically explore and compare the performance of several advanced models on spatial tasks. The dataset includes twelve distinct task types, such as spatial understanding and simple route planning, each with verified and accurate answers. We evaluated multiple models, including OpenAI's gpt-3.5-turbo, gpt-4-turbo, gpt-4o, ZhipuAI's glm-4, Anthropic's claude-3-sonnet-20240229, and MoonShot's moonshot-v1-8k, using a two-phase testing approach. First, we conducted zero-shot testing. Then, we categorized the dataset by difficulty and performed prompt-tuning tests. Results show that gpt-4o achieved the highest overall accuracy in the first phase, with an average of 71.3%. Although moonshot-v1-8k slightly underperformed overall, it outperformed gpt-4o in place name recognition tasks. The study also highlights the impact of prompt strategies on model performance in specific tasks. For instance, the Chain-of-Thought (CoT) strategy increased gpt-4o's accuracy in simple route planning from 12.4% to 87.5%, while a one-shot strategy improved moonshot-v1-8k's accuracy in mapping tasks from 10.1% to 76.3%.
title Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2408.14438