CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Huan-ang, Zhang, Zikang, Luo, Tianwei, Yang, Kaisen, Juan, Xinzhe, Qiu, Jiahao, Chen, Tianxing, He, Bingxiang, Zhao, Hao, Zhou, Hao, Liu, Shilong, Wang, Mengdi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912799445221376
author Gao, Huan-ang
Zhang, Zikang
Luo, Tianwei
Yang, Kaisen
Juan, Xinzhe
Qiu, Jiahao
Chen, Tianxing
He, Bingxiang
Zhao, Hao
Zhou, Hao
Liu, Shilong
Wang, Mengdi
author_facet Gao, Huan-ang
Zhang, Zikang
Luo, Tianwei
Yang, Kaisen
Juan, Xinzhe
Qiu, Jiahao
Chen, Tianxing
He, Bingxiang
Zhao, Hao
Zhou, Hao
Liu, Shilong
Wang, Mengdi
contents Large Language Model (LLM) agents, while proficient in the digital realm, face a significant gap in physical-world deployment due to the challenge of forming and maintaining a robust spatial mental model. We identify three core cognitive challenges hindering this transition: spatial reasoning, long-horizon state tracking via mental simulation, and active exploration under partial observation. To isolate and evaluate these faculties, we introduce CubeBench, a novel generative benchmark centered on the Rubik's Cube. CubeBench uses a three-tiered diagnostic framework that progressively assesses agent capabilities, from foundational state tracking with full symbolic information to active exploration with only partial visual data. Our experiments on leading LLMs reveal critical limitations, including a uniform 0.00% pass rate on all long-horizon tasks, exposing a fundamental failure in long-term planning. We also propose a diagnostic framework to isolate these cognitive bottlenecks by providing external solver tools. By analyzing the failure modes, we provide key insights to guide the development of more physically-grounded intelligent agents.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23328
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations
Gao, Huan-ang
Zhang, Zikang
Luo, Tianwei
Yang, Kaisen
Juan, Xinzhe
Qiu, Jiahao
Chen, Tianxing
He, Bingxiang
Zhao, Hao
Zhou, Hao
Liu, Shilong
Wang, Mengdi
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Large Language Model (LLM) agents, while proficient in the digital realm, face a significant gap in physical-world deployment due to the challenge of forming and maintaining a robust spatial mental model. We identify three core cognitive challenges hindering this transition: spatial reasoning, long-horizon state tracking via mental simulation, and active exploration under partial observation. To isolate and evaluate these faculties, we introduce CubeBench, a novel generative benchmark centered on the Rubik's Cube. CubeBench uses a three-tiered diagnostic framework that progressively assesses agent capabilities, from foundational state tracking with full symbolic information to active exploration with only partial visual data. Our experiments on leading LLMs reveal critical limitations, including a uniform 0.00% pass rate on all long-horizon tasks, exposing a fundamental failure in long-term planning. We also propose a diagnostic framework to isolate these cognitive bottlenecks by providing external solver tools. By analyzing the failure modes, we provide key insights to guide the development of more physically-grounded intelligent agents.
title CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.23328