Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cohn, Anthony G, Blackwell, Robert E
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2507.12059
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917070126448640
author Cohn, Anthony G
Blackwell, Robert E
author_facet Cohn, Anthony G
Blackwell, Robert E
contents We investigate the abilities of 28 Large language Models (LLMs) to reason about cardinal directions (CDs) using a benchmark generated from a set of templates, extensively testing an LLM's ability to determine the correct CD given a particular scenario. The templates allow for a number of degrees of variation such as means of locomotion of the agent involved, and whether set in the first, second or third person. Even the newer Large Reasoning Models are unable to reliably determine the correct CD for all questions. This paper summarises and extends earlier work presented at COSIT-24.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12059
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
Cohn, Anthony G
Blackwell, Robert E
Computation and Language
We investigate the abilities of 28 Large language Models (LLMs) to reason about cardinal directions (CDs) using a benchmark generated from a set of templates, extensively testing an LLM's ability to determine the correct CD given a particular scenario. The templates allow for a number of degrees of variation such as means of locomotion of the agent involved, and whether set in the first, second or third person. Even the newer Large Reasoning Models are unable to reliably determine the correct CD for all questions. This paper summarises and extends earlier work presented at COSIT-24.
title Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
topic Computation and Language
url https://arxiv.org/abs/2507.12059