Planning in the LLM Era: Building for Reliability and Efficiency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Katz, Michael, Kokel, Harsha, Srinivas, Kavitha, Sohrabi, Shirin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thought of Search: Planning with Language Models Through The Lens of Efficiency
von: Katz, Michael, et al.
Veröffentlicht: (2024)
von: Katz, Michael, et al.
Veröffentlicht: (2024)
Large Language Models as Planning Domain Generators
von: Oswald, James, et al.
Veröffentlicht: (2024)
von: Oswald, James, et al.
Veröffentlicht: (2024)
ACPBench: Reasoning about Action, Change, and Planning
von: Kokel, Harsha, et al.
Veröffentlicht: (2024)
von: Kokel, Harsha, et al.
Veröffentlicht: (2024)
ACPBench Hard: Unrestrained Reasoning about Action, Change, and Planning
von: Kokel, Harsha, et al.
Veröffentlicht: (2025)
von: Kokel, Harsha, et al.
Veröffentlicht: (2025)
Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents
von: Sohrabi, Shirin, et al.
Veröffentlicht: (2026)
von: Sohrabi, Shirin, et al.
Veröffentlicht: (2026)
Automating Thought of Search: A Journey Towards Soundness and Completeness
von: Cao, Daniel, et al.
Veröffentlicht: (2024)
von: Cao, Daniel, et al.
Veröffentlicht: (2024)
Model Space Reasoning as Search in Feedback Space for Planning Domain Generation
von: Oswald, James, et al.
Veröffentlicht: (2026)
von: Oswald, James, et al.
Veröffentlicht: (2026)
Make Planning Research Rigorous Again!
von: Katz, Michael, et al.
Veröffentlicht: (2025)
von: Katz, Michael, et al.
Veröffentlicht: (2025)
QueryGym: Step-by-Step Interaction with Relational Databases
von: Ananthakrishnan, Haritha, et al.
Veröffentlicht: (2025)
von: Ananthakrishnan, Haritha, et al.
Veröffentlicht: (2025)
Seemingly Simple Planning Problems are Computationally Challenging: The Countdown Game
von: Katz, Michael, et al.
Veröffentlicht: (2025)
von: Katz, Michael, et al.
Veröffentlicht: (2025)
Unifying and Certifying Top-Quality Planning
von: Katz, Michael, et al.
Veröffentlicht: (2024)
von: Katz, Michael, et al.
Veröffentlicht: (2024)
Can Cross Encoders Produce Useful Sentence Embeddings?
von: Ananthakrishnan, Haritha, et al.
Veröffentlicht: (2025)
von: Ananthakrishnan, Haritha, et al.
Veröffentlicht: (2025)
Some Orders Are Important: Partially Preserving Orders in Top-Quality Planning
von: Katz, Michael, et al.
Veröffentlicht: (2024)
von: Katz, Michael, et al.
Veröffentlicht: (2024)
Less is More: Learning Graph Tasks with Just LLMs
von: Shirai, Sola, et al.
Veröffentlicht: (2025)
von: Shirai, Sola, et al.
Veröffentlicht: (2025)
TravelBench : Exploring LLM Performance in Low-Resource Domains
von: Billa, Srinivas, et al.
Veröffentlicht: (2025)
von: Billa, Srinivas, et al.
Veröffentlicht: (2025)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
von: Zhao, Jingqian, et al.
Veröffentlicht: (2025)
von: Zhao, Jingqian, et al.
Veröffentlicht: (2025)
Improved Generalized Planning with LLMs through Strategy Refinement and Reflection
von: Stein, Katharina, et al.
Veröffentlicht: (2025)
von: Stein, Katharina, et al.
Veröffentlicht: (2025)
MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
von: Phan, Peter, et al.
Veröffentlicht: (2025)
von: Phan, Peter, et al.
Veröffentlicht: (2025)
Unified Mind Model: Reimagining Autonomous Agents in the LLM Era
von: Hu, Pengbo, et al.
Veröffentlicht: (2025)
von: Hu, Pengbo, et al.
Veröffentlicht: (2025)
LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
von: Pihulski, Dzmitry, et al.
Veröffentlicht: (2025)
von: Pihulski, Dzmitry, et al.
Veröffentlicht: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
von: Gonzalez, Miguel Angel Alvarado, et al.
Veröffentlicht: (2025)
von: Gonzalez, Miguel Angel Alvarado, et al.
Veröffentlicht: (2025)
Reliable and diverse evaluation of LLM medical knowledge mastery
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2024)
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
An LLM Maturity Model for Reliable and Transparent Text-to-Query
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
OckBench: Measuring the Efficiency of LLM Reasoning
von: Du, Zheng, et al.
Veröffentlicht: (2025)
von: Du, Zheng, et al.
Veröffentlicht: (2025)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
Extending Minimal Pairs with Ordinal Surprisal Curves and Entropy Across Applied Domains
von: Katz, Andrew
Veröffentlicht: (2026)
von: Katz, Andrew
Veröffentlicht: (2026)
Robust Planning with Compound LLM Architectures: An LLM-Modulo Approach
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
von: Wei, Hui, et al.
Veröffentlicht: (2025)
von: Wei, Hui, et al.
Veröffentlicht: (2025)
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Improving LLM Reliability with RAG in Religious Question-Answering: MufassirQAS
von: Alan, Ahmet Yusuf, et al.
Veröffentlicht: (2024)
von: Alan, Ahmet Yusuf, et al.
Veröffentlicht: (2024)
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
von: Dou, Zhihao, et al.
Veröffentlicht: (2025)
von: Dou, Zhihao, et al.
Veröffentlicht: (2025)
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
von: Utami, Nabelanita, et al.
Veröffentlicht: (2026)
von: Utami, Nabelanita, et al.
Veröffentlicht: (2026)
Towards Reliable and Practical LLM Security Evaluations via Bayesian Modelling
von: Llewellyn, Mary, et al.
Veröffentlicht: (2025)
von: Llewellyn, Mary, et al.
Veröffentlicht: (2025)
The Silent Vote: Improving Zero-Shot LLM Reliability by Aggregating Semantic Neighborhoods
von: Badhe, Sanket, et al.
Veröffentlicht: (2026)
von: Badhe, Sanket, et al.
Veröffentlicht: (2026)
Experiences Build Characters: The Linguistic Origins and Functional Impact of LLM Personality
von: Wang, Xi, et al.
Veröffentlicht: (2026)
von: Wang, Xi, et al.
Veröffentlicht: (2026)
LLM-based Prompt Ensemble for Reliable Medical Entity Recognition from EHRs
von: Islam, K M Sajjadul, et al.
Veröffentlicht: (2025)
von: Islam, K M Sajjadul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Thought of Search: Planning with Language Models Through The Lens of Efficiency
von: Katz, Michael, et al.
Veröffentlicht: (2024) -
Large Language Models as Planning Domain Generators
von: Oswald, James, et al.
Veröffentlicht: (2024) -
ACPBench: Reasoning about Action, Change, and Planning
von: Kokel, Harsha, et al.
Veröffentlicht: (2024) -
ACPBench Hard: Unrestrained Reasoning about Action, Change, and Planning
von: Kokel, Harsha, et al.
Veröffentlicht: (2025) -
Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents
von: Sohrabi, Shirin, et al.
Veröffentlicht: (2026)