Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Fangjun, Hogg, David C., Cohn, Anthony G. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
by: Li, Fangjun, et al.
Published: (2024)
by: Li, Fangjun, et al.
Published: (2024)
A Prolog Program for Bottom-up Evaluation
by: Warren, David S.
Published: (2025)
by: Warren, David S.
Published: (2025)
SpotIt+: Verification-based Text-to-SQL Evaluation with Database Constraints
by: Tremante, Andrew, et al.
Published: (2026)
by: Tremante, Andrew, et al.
Published: (2026)
Efficient Evaluation of Arbitrary Relational Calculus Queries
by: Raszyk, Martin, et al.
Published: (2022)
by: Raszyk, Martin, et al.
Published: (2022)
From Time to Space: The Impact of Linearity in Higher-Order Datalog
by: Charalambidis, Angelos, et al.
Published: (2026)
by: Charalambidis, Angelos, et al.
Published: (2026)
The Power of Negation in Higher-Order Datalog
by: Charalambidis, Angelos, et al.
Published: (2025)
by: Charalambidis, Angelos, et al.
Published: (2025)
Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing
by: Świechowski, Maciej, et al.
Published: (2026)
by: Świechowski, Maciej, et al.
Published: (2026)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
by: Chen, Michael K., et al.
Published: (2025)
by: Chen, Michael K., et al.
Published: (2025)
On The Reasonable Effectiveness of Relational Diagrams: Explaining Relational Query Patterns and the Pattern Expressiveness of Relational Languages
by: Gatterbauer, Wolfgang, et al.
Published: (2024)
by: Gatterbauer, Wolfgang, et al.
Published: (2024)
Work-Efficient Query Evaluation in Constant Time with PRAMs
by: Keppeler, Jens, et al.
Published: (2023)
by: Keppeler, Jens, et al.
Published: (2023)
Trade-offs in Static and Dynamic Evaluation of Hierarchical Queries
by: Kara, Ahmet, et al.
Published: (2019)
by: Kara, Ahmet, et al.
Published: (2019)
SpotIt: Evaluating Text-to-SQL Evaluation with Formal Verification
by: Klopfenstein, Rocky, et al.
Published: (2025)
by: Klopfenstein, Rocky, et al.
Published: (2025)
Common Foundations for Recursive Shape Languages
by: Ahmetaj, Shqiponja, et al.
Published: (2026)
by: Ahmetaj, Shqiponja, et al.
Published: (2026)
PM-LLM-Benchmark: Evaluating Large Language Models on Process Mining Tasks
by: Berti, Alessandro, et al.
Published: (2024)
by: Berti, Alessandro, et al.
Published: (2024)
Completeness of Relational Algebra via Cylindric Algebra
by: Laštovička, Jan
Published: (2026)
by: Laštovička, Jan
Published: (2026)
Structural Indexing of Relational Databases for the Evaluation of Free-Connex Acyclic Conjunctive Queries
by: Riveros, Cristian, et al.
Published: (2026)
by: Riveros, Cristian, et al.
Published: (2026)
Database Research needs an Abstract Relational Query Language
by: Gatterbauer, Wolfgang, et al.
Published: (2025)
by: Gatterbauer, Wolfgang, et al.
Published: (2025)
Using Color Refinement to Boost Enumeration and Counting for Acyclic CQs of Binary Schemas
by: Riveros, Cristian, et al.
Published: (2024)
by: Riveros, Cristian, et al.
Published: (2024)
Logical Foundations and Complexity of 4QL, a Query Language with Unrestricted Negation
by: Maluszynski, Jan, et al.
Published: (2010)
by: Maluszynski, Jan, et al.
Published: (2010)
Using Large Language Models for (De-)Formalization and Natural Argumentation Exercises for Beginner's Students
by: Carl, Merlin
Published: (2023)
by: Carl, Merlin
Published: (2023)
Generics and Default Reasoning in Large Language Models
by: Kirkpatrick, James Ravi, et al.
Published: (2025)
by: Kirkpatrick, James Ravi, et al.
Published: (2025)
Homomorphism Problems in Graph Databases and Automatic Structures
by: Morvan, Rémi
Published: (2025)
by: Morvan, Rémi
Published: (2025)
Complex event recognition under time constraints: towards a formal framework for efficient query evaluation
by: García, Julián, et al.
Published: (2025)
by: García, Julián, et al.
Published: (2025)
FC-Datalog as a Framework for Efficient String Querying
by: Bell, Owen M., et al.
Published: (2025)
by: Bell, Owen M., et al.
Published: (2025)
A Trichotomy for Regular Trail Queries
by: Martens, Wim, et al.
Published: (2019)
by: Martens, Wim, et al.
Published: (2019)
A formal query language and automata model for aggregation in complex event recognition
by: Bourhis, Pierre, et al.
Published: (2026)
by: Bourhis, Pierre, et al.
Published: (2026)
Semantic Tree-Width and Path-Width of Conjunctive Regular Path Queries
by: Figueira, Diego, et al.
Published: (2022)
by: Figueira, Diego, et al.
Published: (2022)
Characterizing Data Dependencies Then and Now
by: Kolaitis, Phokion G., et al.
Published: (2024)
by: Kolaitis, Phokion G., et al.
Published: (2024)
Restricted Chase Termination: You Want More than Fairness
by: Carral, David, et al.
Published: (2025)
by: Carral, David, et al.
Published: (2025)
Disjunctions of Two Dependence Atoms
by: Fröhlich, Nicolas, et al.
Published: (2025)
by: Fröhlich, Nicolas, et al.
Published: (2025)
A Datalog Framework for Conflict-Free Replicated Data Types
by: Yanakieva, Elena, et al.
Published: (2026)
by: Yanakieva, Elena, et al.
Published: (2026)
Query Languages for Machine-Learning Models
by: Grohe, Martin
Published: (2026)
by: Grohe, Martin
Published: (2026)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
by: Cohn, Anthony G, et al.
Published: (2024)
by: Cohn, Anthony G, et al.
Published: (2024)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
by: Cohn, Anthony G, et al.
Published: (2025)
by: Cohn, Anthony G, et al.
Published: (2025)
Pushing the Limit: Verified Performance-Optimal Causally-Consistent Database Transactions
by: Ghasemirad, Shabnam, et al.
Published: (2024)
by: Ghasemirad, Shabnam, et al.
Published: (2024)
From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
by: Bayat, Farima Fatahi, et al.
Published: (2025)
by: Bayat, Farima Fatahi, et al.
Published: (2025)
Queries With Exact Truth Values in Paraconsistent Description Logics
by: Bienvenu, Meghyn, et al.
Published: (2024)
by: Bienvenu, Meghyn, et al.
Published: (2024)
Constructive Interpolation and Concept-Based Beth Definability for Description Logics via Sequents
by: Lyon, Tim S., et al.
Published: (2024)
by: Lyon, Tim S., et al.
Published: (2024)
Properties for Paths in Graph Databases
by: Orejas, Fernando, et al.
Published: (2025)
by: Orejas, Fernando, et al.
Published: (2025)
Similar Items
-
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
by: Li, Fangjun, et al.
Published: (2024) -
A Prolog Program for Bottom-up Evaluation
by: Warren, David S.
Published: (2025) -
SpotIt+: Verification-based Text-to-SQL Evaluation with Database Constraints
by: Tremante, Andrew, et al.
Published: (2026) -
Efficient Evaluation of Arbitrary Relational Calculus Queries
by: Raszyk, Martin, et al.
Published: (2022) -
From Time to Space: The Impact of Linearity in Higher-Order Datalog
by: Charalambidis, Angelos, et al.
Published: (2026)