AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Weiyi, Chen, Xinchi, Gong, Jingjing, Huang, Xuanjing, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
by: Fei, Zhaoye, et al.
Published: (2025)
by: Fei, Zhaoye, et al.
Published: (2025)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
by: Xu, Ningyu, et al.
Published: (2026)
by: Xu, Ningyu, et al.
Published: (2026)
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
by: Men, Tianyi, et al.
Published: (2025)
by: Men, Tianyi, et al.
Published: (2025)
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
by: Parmar, Mihir, et al.
Published: (2025)
by: Parmar, Mihir, et al.
Published: (2025)
UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models
by: Zheng, Yu, et al.
Published: (2025)
by: Zheng, Yu, et al.
Published: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
by: Wang, Yubang, et al.
Published: (2026)
by: Wang, Yubang, et al.
Published: (2026)
Revealing emergent human-like conceptual representations from language prediction
by: Xu, Ningyu, et al.
Published: (2025)
by: Xu, Ningyu, et al.
Published: (2025)
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
by: Shi, Junhao, et al.
Published: (2025)
by: Shi, Junhao, et al.
Published: (2025)
Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems
by: Wang, Ling, et al.
Published: (2026)
by: Wang, Ling, et al.
Published: (2026)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
by: Zhang, Yinger, et al.
Published: (2026)
by: Zhang, Yinger, et al.
Published: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026)
by: Yang, Jie, et al.
Published: (2026)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
by: Valmeekam, Karthik, et al.
Published: (2024)
by: Valmeekam, Karthik, et al.
Published: (2024)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
by: Saha, Swarnadeep, et al.
Published: (2025)
by: Saha, Swarnadeep, et al.
Published: (2025)
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
by: Ding, Deming, et al.
Published: (2026)
by: Ding, Deming, et al.
Published: (2026)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
by: Wang, Peng, et al.
Published: (2026)
by: Wang, Peng, et al.
Published: (2026)
General Agentic Planning Through Simulative Reasoning with World Models
by: Deng, Mingkai, et al.
Published: (2025)
by: Deng, Mingkai, et al.
Published: (2025)
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning
by: Deng, Mingkai, et al.
Published: (2026)
by: Deng, Mingkai, et al.
Published: (2026)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
by: Bu, Yuyan, et al.
Published: (2026)
by: Bu, Yuyan, et al.
Published: (2026)
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
by: Dou, Zhihao, et al.
Published: (2025)
by: Dou, Zhihao, et al.
Published: (2025)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
by: Choukrani, Omar, et al.
Published: (2025)
by: Choukrani, Omar, et al.
Published: (2025)
Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning
by: Fouilhé, Guilhem, et al.
Published: (2026)
by: Fouilhé, Guilhem, et al.
Published: (2026)
CocoaBench: Evaluating Unified Digital Agents in the Wild
by: CocoaBench Team, et al.
Published: (2026)
by: CocoaBench Team, et al.
Published: (2026)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
RECAP: REwriting Conversations for Intent Understanding in Agentic Planning
by: Mitra, Kushan, et al.
Published: (2025)
by: Mitra, Kushan, et al.
Published: (2025)
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
by: Zong, Yi, et al.
Published: (2024)
by: Zong, Yi, et al.
Published: (2024)
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
Self-Polish: Enhance Reasoning in Large Language Models via Problem Refinement
by: Xi, Zhiheng, et al.
Published: (2023)
by: Xi, Zhiheng, et al.
Published: (2023)
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
by: Yin, Zhangyue, et al.
Published: (2025)
by: Yin, Zhangyue, et al.
Published: (2025)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
Matrix as Plan: Structured Logical Reasoning with Feedback-Driven Replanning
by: Chen, Ke, et al.
Published: (2026)
by: Chen, Ke, et al.
Published: (2026)
Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
by: Li, Mufei, et al.
Published: (2025)
by: Li, Mufei, et al.
Published: (2025)
Routine: A Structural Planning Framework for LLM Agent System in Enterprise
by: Zeng, Guancheng, et al.
Published: (2025)
by: Zeng, Guancheng, et al.
Published: (2025)
Reinforced Context Order Recovery for Adaptive Reasoning and Planning
by: Ma, Long, et al.
Published: (2025)
by: Ma, Long, et al.
Published: (2025)
InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding
by: Jiayang, Cheng, et al.
Published: (2025)
by: Jiayang, Cheng, et al.
Published: (2025)
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
by: Xi, Zhiheng, et al.
Published: (2024)
by: Xi, Zhiheng, et al.
Published: (2024)
Iterative Formalization and Planning in Partially Observable Environments
by: Gong, Liancheng, et al.
Published: (2025)
by: Gong, Liancheng, et al.
Published: (2025)
Similar Items
-
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
by: Fei, Zhaoye, et al.
Published: (2025) -
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025) -
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
by: Xu, Ningyu, et al.
Published: (2026) -
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
by: Men, Tianyi, et al.
Published: (2025) -
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
by: Parmar, Mihir, et al.
Published: (2025)