ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Rui, Lu, Dakuan, Zhao, Zicheng, Tan, Xiaoyu, Wang, Xintao, Yuan, Siyu, Chen, Jiangjie, Xu, Yinghui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MINDECHO: Role-Playing Language Agents for Key Opinion Leaders
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain
by: Lu, Dakuan, et al.
Published: (2025)
by: Lu, Dakuan, et al.
Published: (2025)
EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
by: Yuan, Siyu, et al.
Published: (2024)
by: Yuan, Siyu, et al.
Published: (2024)
Character is Destiny: Can Role-Playing Language Agents Make Persona-Driven Decisions?
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
CAD-GPT: Synthesising CAD Construction Sequence with Spatial Reasoning-Enhanced Multimodal LLMs
by: Wang, Siyu, et al.
Published: (2024)
by: Wang, Siyu, et al.
Published: (2024)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
by: Li, Yinghui, et al.
Published: (2025)
by: Li, Yinghui, et al.
Published: (2025)
Struct-X: Enhancing Large Language Models Reasoning with Structured Data
by: Tan, Xiaoyu, et al.
Published: (2024)
by: Tan, Xiaoyu, et al.
Published: (2024)
Towards Collaborative Intelligence: Propagating Intentions and Reasoning for Multi-Agent Coordination with Large Language Models
by: Qiu, Xihe, et al.
Published: (2024)
by: Qiu, Xihe, et al.
Published: (2024)
ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge Base
by: Yuan, Siyu, et al.
Published: (2023)
by: Yuan, Siyu, et al.
Published: (2023)
The Law of Multi-Model Collaboration: Scaling Limits of Model Ensembling for Large Language Models
by: Lu, Dakuan, et al.
Published: (2025)
by: Lu, Dakuan, et al.
Published: (2025)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
Reasoning Fails Where Step Flow Breaks
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
by: Shi, Shaojie, et al.
Published: (2026)
by: Shi, Shaojie, et al.
Published: (2026)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
by: Wang, Xintao, et al.
Published: (2025)
by: Wang, Xintao, et al.
Published: (2025)
VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
by: Yao, Jian, et al.
Published: (2025)
by: Yao, Jian, et al.
Published: (2025)
SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
by: Zheng, Xiang, et al.
Published: (2026)
by: Zheng, Xiang, et al.
Published: (2026)
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Driven Prolog-based Chain-of-Thought
by: Tan, Xiaoyu, et al.
Published: (2024)
by: Tan, Xiaoyu, et al.
Published: (2024)
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
by: Cheng, Ziming, et al.
Published: (2025)
by: Cheng, Ziming, et al.
Published: (2025)
From Building Blocks to Planning: Multi-Step Spatial Reasoning in LLMs with Reinforcement Learning
by: Tahmasbi, Amir, et al.
Published: (2025)
by: Tahmasbi, Amir, et al.
Published: (2025)
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
by: Du, Yongkang, et al.
Published: (2026)
by: Du, Yongkang, et al.
Published: (2026)
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
by: Liu, Chengwu, et al.
Published: (2025)
by: Liu, Chengwu, et al.
Published: (2025)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
by: Yang, Tianyu, et al.
Published: (2026)
by: Yang, Tianyu, et al.
Published: (2026)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
by: Cao, Lang, et al.
Published: (2024)
by: Cao, Lang, et al.
Published: (2024)
Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works
by: Yuan, Xinfeng, et al.
Published: (2024)
by: Yuan, Xinfeng, et al.
Published: (2024)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
by: Hong, Zijin, et al.
Published: (2025)
by: Hong, Zijin, et al.
Published: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
by: Tang, Kexian, et al.
Published: (2025)
by: Tang, Kexian, et al.
Published: (2025)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
by: Chen, Jiangjie, et al.
Published: (2023)
by: Chen, Jiangjie, et al.
Published: (2023)
FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
by: Pyo, Jiyoon, et al.
Published: (2025)
by: Pyo, Jiyoon, et al.
Published: (2025)
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
by: Du, Lingxiao, et al.
Published: (2025)
by: Du, Lingxiao, et al.
Published: (2025)
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
by: Lin, Yen-Ting, et al.
Published: (2025)
by: Lin, Yen-Ting, et al.
Published: (2025)
Step-level Value Preference Optimization for Mathematical Reasoning
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic
by: Xu, Chuou, et al.
Published: (2026)
by: Xu, Chuou, et al.
Published: (2026)
Similar Items
-
MINDECHO: Role-Playing Language Agents for Key Opinion Leaders
by: Xu, Rui, et al.
Published: (2024) -
Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents
by: Xu, Rui, et al.
Published: (2025) -
SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain
by: Lu, Dakuan, et al.
Published: (2025) -
EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
by: Yuan, Siyu, et al.
Published: (2024) -
Character is Destiny: Can Role-Playing Language Agents Make Persona-Driven Decisions?
by: Xu, Rui, et al.
Published: (2024)