AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Weiyi, Chen, Xinchi, Gong, Jingjing, Huang, Xuanjing, Qiu, Xipeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908770714517504
author Wang, Weiyi
Chen, Xinchi
Gong, Jingjing
Huang, Xuanjing
Qiu, Xipeng
author_facet Wang, Weiyi
Chen, Xinchi
Gong, Jingjing
Huang, Xuanjing
Qiu, Xipeng
contents Recent advances in agentic Large Language Models (LLMs) have positioned them as generalist planners capable of reasoning and acting across diverse tasks. However, existing agent benchmarks largely focus on symbolic or weakly grounded environments, leaving their performance in physics-constrained real-world domains underexplored. We introduce AstroReason-Bench, a comprehensive benchmark for evaluating agentic planning in Space Planning Problems (SPP), a family of high-stakes problems with heterogeneous objectives, strict physical constraints, and long-horizon decision-making. AstroReason-Bench integrates multiple scheduling regimes, including ground station communication and agile Earth observation, and provides a unified agent-oriented interaction protocol. Evaluating on a range of state-of-the-art open- and closed-source agentic LLM systems, we find that current agents substantially underperform specialized solvers, highlighting key limitations of generalist planning under realistic constraints. AstroReason-Bench offers a challenging and diagnostic testbed for future agentic research.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11354
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
Wang, Weiyi
Chen, Xinchi
Gong, Jingjing
Huang, Xuanjing
Qiu, Xipeng
Artificial Intelligence
Computation and Language
Recent advances in agentic Large Language Models (LLMs) have positioned them as generalist planners capable of reasoning and acting across diverse tasks. However, existing agent benchmarks largely focus on symbolic or weakly grounded environments, leaving their performance in physics-constrained real-world domains underexplored. We introduce AstroReason-Bench, a comprehensive benchmark for evaluating agentic planning in Space Planning Problems (SPP), a family of high-stakes problems with heterogeneous objectives, strict physical constraints, and long-horizon decision-making. AstroReason-Bench integrates multiple scheduling regimes, including ground station communication and agile Earth observation, and provides a unified agent-oriented interaction protocol. Evaluating on a range of state-of-the-art open- and closed-source agentic LLM systems, we find that current agents substantially underperform specialized solvers, highlighting key limitations of generalist planning under realistic constraints. AstroReason-Bench offers a challenging and diagnostic testbed for future agentic research.
title AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.11354