DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yinger, Jiang, Shutong, Li, Renhao, Tu, Jianhong, Su, Yang, Deng, Lianghao, Guo, Xudong, Lv, Chenxu, Lin, Junyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
by: Xiao, Zikai, et al.
Published: (2025)
by: Xiao, Zikai, et al.
Published: (2025)
ToolRM: Towards Agentic Tool-Use Reward Modeling
by: Li, Renhao, et al.
Published: (2025)
by: Li, Renhao, et al.
Published: (2025)
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
by: Hu, Xiaomeng, et al.
Published: (2026)
by: Hu, Xiaomeng, et al.
Published: (2026)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
Table-as-Search: Formulate Long-Horizon Agentic Information Seeking as Table Completion
by: Lan, Tian, et al.
Published: (2026)
by: Lan, Tian, et al.
Published: (2026)
Scaling Agentic Verifier for Competitive Coding
by: Ma, Zeyao, et al.
Published: (2026)
by: Ma, Zeyao, et al.
Published: (2026)
Open Grounded Planning: Challenges and Benchmark Construction
by: Guo, Shiguang, et al.
Published: (2024)
by: Guo, Shiguang, et al.
Published: (2024)
Planning Transformer: Long-Horizon Offline Reinforcement Learning with Planning Tokens
by: Clinton, Joseph, et al.
Published: (2024)
by: Clinton, Joseph, et al.
Published: (2024)
WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning
by: Yu, Xinmiao, et al.
Published: (2026)
by: Yu, Xinmiao, et al.
Published: (2026)
Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
by: Cui, Sijia, et al.
Published: (2025)
by: Cui, Sijia, et al.
Published: (2025)
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
by: Liang, Xiwen, et al.
Published: (2025)
by: Liang, Xiwen, et al.
Published: (2025)
Language Models can Self-Lengthen to Generate Long Texts
by: Quan, Shanghaoran, et al.
Published: (2024)
by: Quan, Shanghaoran, et al.
Published: (2024)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
by: Lee, Yoonsang, et al.
Published: (2026)
by: Lee, Yoonsang, et al.
Published: (2026)
DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
by: Hu, Xavier, et al.
Published: (2026)
by: Hu, Xavier, et al.
Published: (2026)
Reverse Chain: A Generic-Rule for LLMs to Master Multi-API Planning
by: Zhang, Yinger, et al.
Published: (2023)
by: Zhang, Yinger, et al.
Published: (2023)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
by: He, Muyu, et al.
Published: (2026)
by: He, Muyu, et al.
Published: (2026)
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
by: Chen, Pei-An, et al.
Published: (2026)
by: Chen, Pei-An, et al.
Published: (2026)
On the Ability of Transformers to Verify Plans
by: Sarrof, Yash, et al.
Published: (2026)
by: Sarrof, Yash, et al.
Published: (2026)
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
by: Wang, Zehong, et al.
Published: (2026)
by: Wang, Zehong, et al.
Published: (2026)
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
by: Yang, Dongjie, et al.
Published: (2025)
by: Yang, Dongjie, et al.
Published: (2025)
Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models
by: Palnitkar, Aadi, et al.
Published: (2026)
by: Palnitkar, Aadi, et al.
Published: (2026)
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
by: Wang, Zexuan, et al.
Published: (2026)
by: Wang, Zexuan, et al.
Published: (2026)
HiMAP-Travel: Hierarchical Multi-Agent Planning for Long-Horizon Constrained Travel
by: Bui, The Viet, et al.
Published: (2026)
by: Bui, The Viet, et al.
Published: (2026)
Long-Horizon Plan Execution in Large Tool Spaces through Entropy-Guided Branching
by: Wei, Rongzhe, et al.
Published: (2026)
by: Wei, Rongzhe, et al.
Published: (2026)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
General Agentic Planning Through Simulative Reasoning with World Models
by: Deng, Mingkai, et al.
Published: (2025)
by: Deng, Mingkai, et al.
Published: (2025)
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
by: Yen, Howard, et al.
Published: (2025)
by: Yen, Howard, et al.
Published: (2025)
ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities
by: Su, Ying, et al.
Published: (2024)
by: Su, Ying, et al.
Published: (2024)
Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue
by: Lin, Jingjie, et al.
Published: (2026)
by: Lin, Jingjie, et al.
Published: (2026)
Towards Full Delegation: Designing Ideal Agentic Behaviors for Travel Planning
by: Jiang, Song, et al.
Published: (2024)
by: Jiang, Song, et al.
Published: (2024)
LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
by: Mantenoglou, Periklis, et al.
Published: (2025)
by: Mantenoglou, Periklis, et al.
Published: (2025)
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
by: Luo, Haotian, et al.
Published: (2025)
by: Luo, Haotian, et al.
Published: (2025)
ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents
by: Shao, Jie-Jing, et al.
Published: (2024)
by: Shao, Jie-Jing, et al.
Published: (2024)
Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
by: Chen, Qianben, et al.
Published: (2026)
by: Chen, Qianben, et al.
Published: (2026)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
by: Otani, Naoki, et al.
Published: (2026)
by: Otani, Naoki, et al.
Published: (2026)
Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
by: Lumer, Elias, et al.
Published: (2026)
by: Lumer, Elias, et al.
Published: (2026)
MemWeaver: Weaving Hybrid Memories for Traceable Long-Horizon Agentic Reasoning
by: Ye, Juexiang, et al.
Published: (2026)
by: Ye, Juexiang, et al.
Published: (2026)
Similar Items
-
LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
by: Xiao, Zikai, et al.
Published: (2025) -
ToolRM: Towards Agentic Tool-Use Reward Modeling
by: Li, Renhao, et al.
Published: (2025) -
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
by: Hu, Xiaomeng, et al.
Published: (2026) -
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
by: Erdogan, Lutfi Eren, et al.
Published: (2025) -
Table-as-Search: Formulate Long-Horizon Agentic Information Seeking as Table Completion
by: Lan, Tian, et al.
Published: (2026)