QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cohn, Anthony G., Blackwell, Robert E. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
von: Cohn, Anthony G, et al.
Veröffentlicht: (2025)
von: Cohn, Anthony G, et al.
Veröffentlicht: (2025)
Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
von: Li, Fangjun, et al.
Veröffentlicht: (2024)
Can Large Language Models Reason about the Region Connection Calculus?
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
von: McPheat, Lachlan, et al.
Veröffentlicht: (2025)
von: McPheat, Lachlan, et al.
Veröffentlicht: (2025)
QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
von: Puyin, Li, et al.
Veröffentlicht: (2025)
von: Puyin, Li, et al.
Veröffentlicht: (2025)
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
Evaluating the Logical Reasoning Abilities of Large Reasoning Models
von: Liu, Hanmeng, et al.
Veröffentlicht: (2025)
von: Liu, Hanmeng, et al.
Veröffentlicht: (2025)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
von: Zhang, Qiran, et al.
Veröffentlicht: (2026)
von: Zhang, Qiran, et al.
Veröffentlicht: (2026)
Exploring Spatial Representations in the Historical Lake District Texts with LLM-based Relation Extraction
von: Haris, Erum, et al.
Veröffentlicht: (2024)
von: Haris, Erum, et al.
Veröffentlicht: (2024)
Deconstructing Instruction-Following: A New Benchmark for Granular Evaluation of Large Language Model Instruction Compliance Abilities
von: Purpura, Alberto, et al.
Veröffentlicht: (2026)
von: Purpura, Alberto, et al.
Veröffentlicht: (2026)
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024)
von: Blackwell, Robert E., et al.
Veröffentlicht: (2024)
FEABench: Evaluating Language Models on Multiphysics Reasoning Ability
von: Mudur, Nayantara, et al.
Veröffentlicht: (2025)
von: Mudur, Nayantara, et al.
Veröffentlicht: (2025)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
SPAN: Benchmarking and Improving Cross-Calendar Temporal Reasoning of Large Language Models
von: Miao, Zhongjian, et al.
Veröffentlicht: (2025)
von: Miao, Zhongjian, et al.
Veröffentlicht: (2025)
Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
von: Parmar, Mihir, et al.
Veröffentlicht: (2024)
von: Parmar, Mihir, et al.
Veröffentlicht: (2024)
REAL: Benchmarking Abilities of Large Language Models for Housing Transactions and Services
von: Zhu, Kexin, et al.
Veröffentlicht: (2025)
von: Zhu, Kexin, et al.
Veröffentlicht: (2025)
Graph-enhanced Large Language Models in Asynchronous Plan Reasoning
von: Lin, Fangru, et al.
Veröffentlicht: (2024)
von: Lin, Fangru, et al.
Veröffentlicht: (2024)
Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
von: Liu, Weichen, et al.
Veröffentlicht: (2025)
von: Liu, Weichen, et al.
Veröffentlicht: (2025)
Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
von: Gu, Zhouhong, et al.
Veröffentlicht: (2024)
von: Gu, Zhouhong, et al.
Veröffentlicht: (2024)
AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials
von: Lv, Taoyuze, et al.
Veröffentlicht: (2025)
von: Lv, Taoyuze, et al.
Veröffentlicht: (2025)
On the Reasoning Abilities of Masked Diffusion Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2025)
von: Svete, Anej, et al.
Veröffentlicht: (2025)
Towards Reasoning Ability of Small Language Models
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2025)
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2025)
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
von: Wang, Haochuan, et al.
Veröffentlicht: (2024)
von: Wang, Haochuan, et al.
Veröffentlicht: (2024)
MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models
von: Ding, Peng, et al.
Veröffentlicht: (2024)
von: Ding, Peng, et al.
Veröffentlicht: (2024)
Large Language Models in Numberland: A Quick Test of Their Numerical Reasoning Abilities
von: Rahman, Roussel
Veröffentlicht: (2025)
von: Rahman, Roussel
Veröffentlicht: (2025)
Eliciting Causal Abilities in Large Language Models for Reasoning Tasks
von: Wang, Yajing, et al.
Veröffentlicht: (2024)
von: Wang, Yajing, et al.
Veröffentlicht: (2024)
An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning
von: Kaur, Navdeep, et al.
Veröffentlicht: (2025)
von: Kaur, Navdeep, et al.
Veröffentlicht: (2025)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
von: Patel, Nisarg, et al.
Veröffentlicht: (2024)
von: Patel, Nisarg, et al.
Veröffentlicht: (2024)
CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models
von: Chun, Jon, et al.
Veröffentlicht: (2026)
von: Chun, Jon, et al.
Veröffentlicht: (2026)
Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2026)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
von: Li, Fangjun, et al.
Veröffentlicht: (2024) -
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024) -
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
von: Cohn, Anthony G, et al.
Veröffentlicht: (2025) -
Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
von: Li, Fangjun, et al.
Veröffentlicht: (2024) -
Can Large Language Models Reason about the Region Connection Calculus?
von: Cohn, Anthony G, et al.
Veröffentlicht: (2024)