Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cohn, Anthony G, Blackwell, Robert E |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
Can Large Language Models Reason about the Region Connection Calculus?
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024)
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024)
QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi
von: Cohn, Anthony G., et al.
Veröffentlicht: (2026)
von: Cohn, Anthony G., et al.
Veröffentlicht: (2026)
Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models
von: Yun, Tian, et al.
Veröffentlicht: (2025)
von: Yun, Tian, et al.
Veröffentlicht: (2025)
Disentangling Memory and Reasoning Ability in Large Language Models
von: Jin, Mingyu, et al.
Veröffentlicht: (2024)
von: Jin, Mingyu, et al.
Veröffentlicht: (2024)
Revisiting the Graph Reasoning Ability of Large Language Models: Case Studies in Translation, Connectivity and Shortest Path
von: Dai, Xinnan, et al.
Veröffentlicht: (2024)
von: Dai, Xinnan, et al.
Veröffentlicht: (2024)
AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models
von: Zhu, Qin, et al.
Veröffentlicht: (2025)
von: Zhu, Qin, et al.
Veröffentlicht: (2025)
LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models
von: Agashe, Saaket, et al.
Veröffentlicht: (2023)
von: Agashe, Saaket, et al.
Veröffentlicht: (2023)
Graph-enhanced Large Language Models in Asynchronous Plan Reasoning
von: Lin, Fangru, et al.
Veröffentlicht: (2024)
von: Lin, Fangru, et al.
Veröffentlicht: (2024)
RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation
von: Diallo, Aissatou, et al.
Veröffentlicht: (2025)
von: Diallo, Aissatou, et al.
Veröffentlicht: (2025)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
von: Parmar, Mihir, et al.
Veröffentlicht: (2024)
von: Parmar, Mihir, et al.
Veröffentlicht: (2024)
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
Distilling Reasoning Ability from Large Language Models with Adaptive Thinking
von: Chen, Xiaoshu, et al.
Veröffentlicht: (2024)
von: Chen, Xiaoshu, et al.
Veröffentlicht: (2024)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
von: Wan, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wan, Yuxuan, et al.
Veröffentlicht: (2024)
ICLEval: Evaluating In-Context Learning Ability of Large Language Models
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
Eliciting Causal Abilities in Large Language Models for Reasoning Tasks
von: Wang, Yajing, et al.
Veröffentlicht: (2024)
von: Wang, Yajing, et al.
Veröffentlicht: (2024)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
von: Gu, Zhouhong, et al.
Veröffentlicht: (2024)
von: Gu, Zhouhong, et al.
Veröffentlicht: (2024)
FEABench: Evaluating Language Models on Multiphysics Reasoning Ability
von: Mudur, Nayantara, et al.
Veröffentlicht: (2025)
von: Mudur, Nayantara, et al.
Veröffentlicht: (2025)
Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2025)
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2025)
Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
von: Sakai, Yusuke, et al.
Veröffentlicht: (2025)
von: Sakai, Yusuke, et al.
Veröffentlicht: (2025)
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
von: Lee, Janghwan, et al.
Veröffentlicht: (2024)
von: Lee, Janghwan, et al.
Veröffentlicht: (2024)
Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2026)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2026)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
Reasoning Can Hurt the Inductive Abilities of Large Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
von: Patel, Nisarg, et al.
Veröffentlicht: (2024)
von: Patel, Nisarg, et al.
Veröffentlicht: (2024)
A Survey on Enhancing Causal Reasoning Ability of Large Language Models
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
Role-Conditioned Refusals: Evaluating Access Control Reasoning in Large Language Models
von: Klisura, Đorđe, et al.
Veröffentlicht: (2025)
von: Klisura, Đorđe, et al.
Veröffentlicht: (2025)
Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
von: Lee, Seungpil, et al.
Veröffentlicht: (2024)
von: Lee, Seungpil, et al.
Veröffentlicht: (2024)
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
Code-Driven Inductive Synthesis: Enhancing Reasoning Abilities of Large Language Models with Sequences
von: Chen, Kedi, et al.
Veröffentlicht: (2025)
von: Chen, Kedi, et al.
Veröffentlicht: (2025)
Optimizing Language Model's Reasoning Abilities with Weak Supervision
von: Tong, Yongqi, et al.
Veröffentlicht: (2024)
von: Tong, Yongqi, et al.
Veröffentlicht: (2024)
Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation
von: Hida, Rem, et al.
Veröffentlicht: (2024)
von: Hida, Rem, et al.
Veröffentlicht: (2024)
Can Large Language Models Generalize Procedures Across Representations?
von: Lin, Fangru, et al.
Veröffentlicht: (2026)
von: Lin, Fangru, et al.
Veröffentlicht: (2026)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
von: Wang, Weiqi, et al.
Veröffentlicht: (2024)
von: Wang, Weiqi, et al.
Veröffentlicht: (2024)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning
von: Yu, Yahan, et al.
Veröffentlicht: (2026)
von: Yu, Yahan, et al.
Veröffentlicht: (2026)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024) -
Can Large Language Models Reason about the Region Connection Calculus?
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024) -
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024) -
QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi
von: Cohn, Anthony G., et al.
Veröffentlicht: (2026) -
Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
von: Li, Fangjun, et al.
Veröffentlicht: (2024)