Exploring and Benchmarking the Planning Capabilities of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bohnet, Bernd, Nova, Azade, Parisi, Aaron T, Swersky, Kevin, Goshvadi, Katayoon, Dai, Hanjun, Schuurmans, Dale, Fiedel, Noah, Sedghi, Hanie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Large Language Model Planning with Action Sequence Similarity
by: Zhao, Xinran, et al.
Published: (2025)
by: Zhao, Xinran, et al.
Published: (2025)
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
by: Bohnet, Bernd, et al.
Published: (2025)
by: Bohnet, Bernd, et al.
Published: (2025)
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
Analysis of Optimality of Large Language Models on Planning Problems
by: Bohnet, Bernd, et al.
Published: (2026)
by: Bohnet, Bernd, et al.
Published: (2026)
Autoregressive Large Language Models are Computationally Universal
by: Schuurmans, Dale, et al.
Published: (2024)
by: Schuurmans, Dale, et al.
Published: (2024)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
A Comparative Analysis of LLM Adaptation: SFT, LoRA, and ICL in Data-Scarce Scenarios
by: Bohnet, Bernd, et al.
Published: (2025)
by: Bohnet, Bernd, et al.
Published: (2025)
Training Language Models on the Knowledge Graph: Insights on Hallucinations and Their Detectability
by: Hron, Jiri, et al.
Published: (2024)
by: Hron, Jiri, et al.
Published: (2024)
Large Language Models can Learn Rules
by: Zhu, Zhaocheng, et al.
Published: (2023)
by: Zhu, Zhaocheng, et al.
Published: (2023)
Judging with Confidence: Calibrating Autoraters to Preference Distributions
by: Li, Zhuohang, et al.
Published: (2025)
by: Li, Zhuohang, et al.
Published: (2025)
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
by: Zheng, Huaixiu Steven, et al.
Published: (2024)
by: Zheng, Huaixiu Steven, et al.
Published: (2024)
UQE: A Query Engine for Unstructured Databases
by: Dai, Hanjun, et al.
Published: (2024)
by: Dai, Hanjun, et al.
Published: (2024)
Transfer Learning for Text Diffusion Models
by: Han, Kehang, et al.
Published: (2024)
by: Han, Kehang, et al.
Published: (2024)
Universal computation is intrinsic to language model decoding
by: Lewandowski, Alex, et al.
Published: (2026)
by: Lewandowski, Alex, et al.
Published: (2026)
Honest Students from Untrusted Teachers: Learning an Interpretable Question-Answering Pipeline from a Pretrained Language Model
by: Eisenstein, Jacob, et al.
Published: (2022)
by: Eisenstein, Jacob, et al.
Published: (2022)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
by: Jacovi, Alon, et al.
Published: (2024)
by: Jacovi, Alon, et al.
Published: (2024)
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
by: Pauli, Amalie Brogaard, et al.
Published: (2024)
by: Pauli, Amalie Brogaard, et al.
Published: (2024)
Exploring the Capabilities of Prompted Large Language Models in Educational and Assessment Applications
by: Maity, Subhankar, et al.
Published: (2024)
by: Maity, Subhankar, et al.
Published: (2024)
Many-Shot In-Context Learning
by: Agarwal, Rishabh, et al.
Published: (2024)
by: Agarwal, Rishabh, et al.
Published: (2024)
Exploring the Word Sense Disambiguation Capabilities of Large Language Models
by: Basile, Pierpaolo, et al.
Published: (2025)
by: Basile, Pierpaolo, et al.
Published: (2025)
Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
by: Singh, Avi, et al.
Published: (2023)
by: Singh, Avi, et al.
Published: (2023)
Planning for Success: Exploring LLM Long-term Planning Capabilities in Table Understanding
by: Nguyen, Thi-Nhung, et al.
Published: (2025)
by: Nguyen, Thi-Nhung, et al.
Published: (2025)
Exploring State Tracking Capabilities of Large Language Models
by: Rezaee, Kiamehr, et al.
Published: (2025)
by: Rezaee, Kiamehr, et al.
Published: (2025)
Cooperative Strategic Planning Enhances Reasoning Capabilities in Large Language Models
by: Wang, Danqing, et al.
Published: (2024)
by: Wang, Danqing, et al.
Published: (2024)
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models
by: Ren, Yuanyi, et al.
Published: (2024)
by: Ren, Yuanyi, et al.
Published: (2024)
MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code Generation
by: Dai, Jianbo, et al.
Published: (2024)
by: Dai, Jianbo, et al.
Published: (2024)
Towards Better Instruction Following Retrieval Models
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation
by: Itai, Koki, et al.
Published: (2026)
by: Itai, Koki, et al.
Published: (2026)
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
by: Kwan, Wai-Chung, et al.
Published: (2024)
by: Kwan, Wai-Chung, et al.
Published: (2024)
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
by: Gao, Yiming, et al.
Published: (2025)
by: Gao, Yiming, et al.
Published: (2025)
A Survey on Multi-Turn Interaction Capabilities of Large Language Models
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
by: Liu, Yixin, et al.
Published: (2023)
by: Liu, Yixin, et al.
Published: (2023)
PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models
by: Song, Inpyo, et al.
Published: (2025)
by: Song, Inpyo, et al.
Published: (2025)
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
by: Li, Jianling, et al.
Published: (2025)
by: Li, Jianling, et al.
Published: (2025)
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
by: Yang, Rem, et al.
Published: (2025)
by: Yang, Rem, et al.
Published: (2025)
Large Language Models Naively Recover Ethnicity from Individual Records
by: Dasanaike, Noah
Published: (2026)
by: Dasanaike, Noah
Published: (2026)
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models
by: Jiang, Jiyue, et al.
Published: (2024)
by: Jiang, Jiyue, et al.
Published: (2024)
FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
by: Mateega, Spencer, et al.
Published: (2025)
by: Mateega, Spencer, et al.
Published: (2025)
Disinformation Capabilities of Large Language Models
by: Vykopal, Ivan, et al.
Published: (2023)
by: Vykopal, Ivan, et al.
Published: (2023)
Similar Items
-
Improving Large Language Model Planning with Action Sequence Similarity
by: Zhao, Xinran, et al.
Published: (2025) -
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
by: Bohnet, Bernd, et al.
Published: (2025) -
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
by: Bohnet, Bernd, et al.
Published: (2024) -
Analysis of Optimality of Large Language Models on Planning Problems
by: Bohnet, Bernd, et al.
Published: (2026) -
Autoregressive Large Language Models are Computationally Universal
by: Schuurmans, Dale, et al.
Published: (2024)