NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Huaixiu Steven, Mishra, Swaroop, Zhang, Hugh, Chen, Xinyun, Chen, Minmin, Nova, Azade, Hou, Le, Cheng, Heng-Tze, Le, Quoc V., Chi, Ed H., Zhou, Denny |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
by: Zheng, Huaixiu Steven, et al.
Published: (2023)
by: Zheng, Huaixiu Steven, et al.
Published: (2023)
Self-Discover: Large Language Models Self-Compose Reasoning Structures
by: Zhou, Pei, et al.
Published: (2024)
by: Zhou, Pei, et al.
Published: (2024)
Large Language Models Cannot Self-Correct Reasoning Yet
by: Huang, Jie, et al.
Published: (2023)
by: Huang, Jie, et al.
Published: (2023)
EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
Large Language Models as Optimizers
by: Yang, Chengrun, et al.
Published: (2023)
by: Yang, Chengrun, et al.
Published: (2023)
Symbol tuning improves in-context learning in language models
by: Wei, Jerry, et al.
Published: (2023)
by: Wei, Jerry, et al.
Published: (2023)
Premise Order Matters in Reasoning with Large Language Models
by: Chen, Xinyun, et al.
Published: (2024)
by: Chen, Xinyun, et al.
Published: (2024)
Improving Large Language Model Planning with Action Sequence Similarity
by: Zhao, Xinran, et al.
Published: (2025)
by: Zhao, Xinran, et al.
Published: (2025)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
by: Verma, Pulkit, et al.
Published: (2025)
by: Verma, Pulkit, et al.
Published: (2025)
SearchGym: A Modular Infrastructure for Cross-Platform Benchmarking and Hybrid Search Orchestration
by: Hsu, Jerome Tze-Hou
Published: (2026)
by: Hsu, Jerome Tze-Hou
Published: (2026)
Large Language Models as Analogical Reasoners
by: Yasunaga, Michihiro, et al.
Published: (2023)
by: Yasunaga, Michihiro, et al.
Published: (2023)
Large Language Models as Tool Makers
by: Cai, Tianle, et al.
Published: (2023)
by: Cai, Tianle, et al.
Published: (2023)
Simple synthetic data reduces sycophancy in large language models
by: Wei, Jerry, et al.
Published: (2023)
by: Wei, Jerry, et al.
Published: (2023)
Large Language Models as Data Augmenters for Cold-Start Item Recommendation
by: Wang, Jianling, et al.
Published: (2024)
by: Wang, Jianling, et al.
Published: (2024)
Khuynh hướng giải huyền thoại trong văn xuôi Việt Nam đương đại từ 1986 đến nay
by: Le, Quoc Hieu
Published: (2017)
by: Le, Quoc Hieu
Published: (2017)
Reassessing the Impact of Foreign Direct Investment on Environmental Quality in 112 Countries: A Bayesian Quantile Regression Approach
by: Dinh Le Quoc
Published: (2025)
by: Dinh Le Quoc
Published: (2025)
LLMs' ways of seeing User Personas
by: Panda, Swaroop
Published: (2024)
by: Panda, Swaroop
Published: (2024)
Balancing Bank Profits With Sustainable Development Goals: Examining the Pivotal Role of Financial Stability
by: Nguyen Quoc Huy, et al.
Published: (2025)
by: Nguyen Quoc Huy, et al.
Published: (2025)
Reverse Thinking Makes LLMs Stronger Reasoners
by: Chen, Justin Chih-Yao, et al.
Published: (2024)
by: Chen, Justin Chih-Yao, et al.
Published: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
by: Zhou, Yongchao, et al.
Published: (2024)
by: Zhou, Yongchao, et al.
Published: (2024)
Sustainable performance measurement through digital transformation within the sustainable development framework: The mediating effect of supply chain concentration
by: Le Sun, et al.
Published: (2024)
by: Le Sun, et al.
Published: (2024)
Vector-Valued Gaussian Processes for Approximating Divergence- or Rotation-free Vector Fields
by: Gia, Quoc Thong Le, et al.
Published: (2025)
by: Gia, Quoc Thong Le, et al.
Published: (2025)
Transforming peptide hormone prediction: The role of AI in modern proteomics
by: Nguyen Quoc Khanh Le
Published: (2024)
by: Nguyen Quoc Khanh Le
Published: (2024)
Artificial Intelligence in Proteomics Clinical Applications
by: Nguyen Quoc Khanh Le
Published: (2025)
by: Nguyen Quoc Khanh Le
Published: (2025)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
by: Truong, Kimberly Le, et al.
Published: (2025)
by: Truong, Kimberly Le, et al.
Published: (2025)
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
by: Parmar, Mihir, et al.
Published: (2025)
by: Parmar, Mihir, et al.
Published: (2025)
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
by: Chang, Chen-Chi, et al.
Published: (2024)
by: Chang, Chen-Chi, et al.
Published: (2024)
GRAFT: Graph-Tokenized LLMs for Tool Planning
by: Gao, Xinyi, et al.
Published: (2026)
by: Gao, Xinyi, et al.
Published: (2026)
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
by: Bohnet, Bernd, et al.
Published: (2025)
by: Bohnet, Bernd, et al.
Published: (2025)
AGFA-Net: Attention-Guided and Feature-Aggregated Network for Coronary Artery Segmentation using Computed Tomography Angiography
by: Liu, Xinyun, et al.
Published: (2024)
by: Liu, Xinyun, et al.
Published: (2024)
When Machine Learning Meets Importance Sampling: A More Efficient Rare Event Estimation Approach
by: Zhao, Ruoning, et al.
Published: (2025)
by: Zhao, Ruoning, et al.
Published: (2025)
Increasing Consumers’ Hypermarket Visit Intention through Cause-Related Marketing: A Perspective from the Theory of Planned Behaviour
by: Kay Tze Hong
Published: (2019)
by: Kay Tze Hong
Published: (2019)
Large Language Models can Learn Rules
by: Zhu, Zhaocheng, et al.
Published: (2023)
by: Zhu, Zhaocheng, et al.
Published: (2023)
Signature in Code Backdoor Detection, how far are we?
by: Le, Quoc Hung, et al.
Published: (2025)
by: Le, Quoc Hung, et al.
Published: (2025)
Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation
by: Xu, Hefei, et al.
Published: (2025)
by: Xu, Hefei, et al.
Published: (2025)
Evolution of time-fractional stochastic hyperbolic diffusion equations on the unit sphere
by: Alodat, Tareq, et al.
Published: (2024)
by: Alodat, Tareq, et al.
Published: (2024)
Multimodal Contextualized Support for Enhancing Video Retrieval System
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026)
by: Le, Van-Truong
Published: (2026)
Similar Items
-
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
by: Zheng, Huaixiu Steven, et al.
Published: (2023) -
Self-Discover: Large Language Models Self-Compose Reasoning Structures
by: Zhou, Pei, et al.
Published: (2024) -
Large Language Models Cannot Self-Correct Reasoning Yet
by: Huang, Jie, et al.
Published: (2023) -
EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
by: Nie, Allen, et al.
Published: (2024) -
Large Language Models as Optimizers
by: Yang, Chengrun, et al.
Published: (2023)