PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Haoming, Chen, Zhaoliang, Zhang, Jonathan, Liu, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives
by: Mercier, Alexandre Le, et al.
Published: (2026)
by: Mercier, Alexandre Le, et al.
Published: (2026)
RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
by: Shen, Chengzhi, et al.
Published: (2026)
by: Shen, Chengzhi, et al.
Published: (2026)
Symphony: A Decentralized Multi-Agent Framework for Scalable Collective Intelligence
by: Wang, Ji, et al.
Published: (2025)
by: Wang, Ji, et al.
Published: (2025)
LASP: Surveying the State-of-the-Art in Large Language Model-Assisted AI Planning
by: Li, Haoming, et al.
Published: (2024)
by: Li, Haoming, et al.
Published: (2024)
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
by: Li, Zhuofeng, et al.
Published: (2025)
by: Li, Zhuofeng, et al.
Published: (2025)
Verification-Aware Planning for Multi-Agent Systems
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
by: Wan, Ziyu, et al.
Published: (2025)
by: Wan, Ziyu, et al.
Published: (2025)
HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research
by: Zhu, Yinghao, et al.
Published: (2025)
by: Zhu, Yinghao, et al.
Published: (2025)
PAACE: A Plan-Aware Automated Agent Context Engineering Framework
by: Yuksel, Kamer Ali
Published: (2025)
by: Yuksel, Kamer Ali
Published: (2025)
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
by: Lin, Huawei, et al.
Published: (2026)
by: Lin, Huawei, et al.
Published: (2026)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
by: Zhu, Yinghao, et al.
Published: (2025)
by: Zhu, Yinghao, et al.
Published: (2025)
Benchmarking Agentic Workflow Generation
by: Qiao, Shuofei, et al.
Published: (2024)
by: Qiao, Shuofei, et al.
Published: (2024)
Harnessing Multi-Agent LLMs for Complex Engineering Problem-Solving: A Framework for Senior Design Projects
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
Unleashing Diverse Thinking Modes in LLMs through Multi-Agent Collaboration
by: He, Zhixuan, et al.
Published: (2025)
by: He, Zhixuan, et al.
Published: (2025)
Memp: Exploring Agent Procedural Memory
by: Fang, Runnan, et al.
Published: (2025)
by: Fang, Runnan, et al.
Published: (2025)
Towards Reliable ML Feature Engineering via Planning in Constrained-Topology of LLM Agents
by: Thakur, Himanshu, et al.
Published: (2026)
by: Thakur, Himanshu, et al.
Published: (2026)
Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMs
by: Jonelagadda, Aneesh, et al.
Published: (2025)
by: Jonelagadda, Aneesh, et al.
Published: (2025)
Composite Learning Units: Generalized Learning Beyond Parameter Updates to Transform LLMs into Adaptive Reasoners
by: Radha, Santosh Kumar, et al.
Published: (2024)
by: Radha, Santosh Kumar, et al.
Published: (2024)
SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data Science
by: Seo, Wonduk, et al.
Published: (2025)
by: Seo, Wonduk, et al.
Published: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026)
by: Chaudhary, Maheep
Published: (2026)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
MASPO: Joint Prompt Optimization for LLM-based Multi-Agent Systems
by: Wang, Zhexuan, et al.
Published: (2026)
by: Wang, Zhexuan, et al.
Published: (2026)
Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
by: Kozlova, Anna, et al.
Published: (2026)
by: Kozlova, Anna, et al.
Published: (2026)
Agent Planning with World Knowledge Model
by: Qiao, Shuofei, et al.
Published: (2024)
by: Qiao, Shuofei, et al.
Published: (2024)
Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integration and Evaluation using Novel Metrics and Dataset
by: Gabriel, Adrian Garret, et al.
Published: (2024)
by: Gabriel, Adrian Garret, et al.
Published: (2024)
Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
by: Zhu, Yuqi, et al.
Published: (2025)
by: Zhu, Yuqi, et al.
Published: (2025)
Can We Predict Before Executing Machine Learning Agents?
by: Zheng, Jingsheng, et al.
Published: (2026)
by: Zheng, Jingsheng, et al.
Published: (2026)
Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation
by: Li, Ed, et al.
Published: (2025)
by: Li, Ed, et al.
Published: (2025)
Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning tasks
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
TrustAgent: Towards Safe and Trustworthy LLM-based Agents
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
by: Fang, Haoyang, et al.
Published: (2025)
by: Fang, Haoyang, et al.
Published: (2025)
Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
by: Xu, Haoming, et al.
Published: (2026)
by: Xu, Haoming, et al.
Published: (2026)
AgentRec: Agent Recommendation Using Sentence Embeddings Aligned to Human Feedback
by: Park, Joshua, et al.
Published: (2025)
by: Park, Joshua, et al.
Published: (2025)
MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
by: Wen, Xiaoyu, et al.
Published: (2026)
by: Wen, Xiaoyu, et al.
Published: (2026)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
by: Liu, Yuhan, et al.
Published: (2024)
by: Liu, Yuhan, et al.
Published: (2024)
DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration
by: Nourzad, Narjes, et al.
Published: (2025)
by: Nourzad, Narjes, et al.
Published: (2025)
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
by: Sarkar, Bidipta, et al.
Published: (2025)
by: Sarkar, Bidipta, et al.
Published: (2025)
Similar Items
-
GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives
by: Mercier, Alexandre Le, et al.
Published: (2026) -
RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
by: Shen, Chengzhi, et al.
Published: (2026) -
Symphony: A Decentralized Multi-Agent Framework for Scalable Collective Intelligence
by: Wang, Ji, et al.
Published: (2025) -
LASP: Surveying the State-of-the-Art in Large Language Model-Assisted AI Planning
by: Li, Haoming, et al.
Published: (2024) -
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
by: Li, Zhuofeng, et al.
Published: (2025)