Saved in:
| Main Authors: | Clinton, Joseph, Lieck, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.09513 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
by: Wang, Zehong, et al.
Published: (2026)
by: Wang, Zehong, et al.
Published: (2026)
LLoCO: Learning Long Contexts Offline
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Guiding Language Model Reasoning with Planning Tokens
by: Wang, Xinyi, et al.
Published: (2023)
by: Wang, Xinyi, et al.
Published: (2023)
Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
On the Ability of Transformers to Verify Plans
by: Sarrof, Yash, et al.
Published: (2026)
by: Sarrof, Yash, et al.
Published: (2026)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
by: Wang, Huaijie, et al.
Published: (2024)
by: Wang, Huaijie, et al.
Published: (2024)
Scaling Optimal LR Across Token Horizons
by: Bjorck, Johan, et al.
Published: (2024)
by: Bjorck, Johan, et al.
Published: (2024)
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
by: Wang, Siwei, et al.
Published: (2025)
by: Wang, Siwei, et al.
Published: (2025)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning
by: Wang, Yinuo, et al.
Published: (2026)
by: Wang, Yinuo, et al.
Published: (2026)
SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation
by: Zhou, Zihan, et al.
Published: (2024)
by: Zhou, Zihan, et al.
Published: (2024)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
by: Pang, Jing-Cheng, et al.
Published: (2024)
by: Pang, Jing-Cheng, et al.
Published: (2024)
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
LHAW: Controllable Underspecification for Long-Horizon Tasks
by: Pu, George, et al.
Published: (2026)
by: Pu, George, et al.
Published: (2026)
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
by: Gur, Izzeddin, et al.
Published: (2023)
by: Gur, Izzeddin, et al.
Published: (2023)
TodoEvolve: Learning to Architect Agent Planning Systems
by: Liu, Jiaxi, et al.
Published: (2026)
by: Liu, Jiaxi, et al.
Published: (2026)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
by: Ge, Tao, et al.
Published: (2026)
by: Ge, Tao, et al.
Published: (2026)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
Conversational Planning for Personal Plans
by: Christakopoulou, Konstantina, et al.
Published: (2025)
by: Christakopoulou, Konstantina, et al.
Published: (2025)
Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning
by: Hwang, Jaebak, et al.
Published: (2025)
by: Hwang, Jaebak, et al.
Published: (2025)
ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models
by: Wang, Siwei, et al.
Published: (2024)
by: Wang, Siwei, et al.
Published: (2024)
PRInTS: Reward Modeling for Long-Horizon Information Seeking
by: Lee, Jaewoo, et al.
Published: (2025)
by: Lee, Jaewoo, et al.
Published: (2025)
Reasoning Planning for Language Models
by: Nguyen, Bao, et al.
Published: (2025)
by: Nguyen, Bao, et al.
Published: (2025)
Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents
by: Zambrano, Alejandra, et al.
Published: (2026)
by: Zambrano, Alejandra, et al.
Published: (2026)
Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models
by: Liu, Youwei, et al.
Published: (2026)
by: Liu, Youwei, et al.
Published: (2026)
AgentFold: Long-Horizon Web Agents with Proactive Context Management
by: Ye, Rui, et al.
Published: (2025)
by: Ye, Rui, et al.
Published: (2025)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
by: Yeo, Woongyeng, et al.
Published: (2026)
by: Yeo, Woongyeng, et al.
Published: (2026)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
by: Zhong, Han, et al.
Published: (2024)
by: Zhong, Han, et al.
Published: (2024)
Explanatory Summarization with Discourse-Driven Planning
by: Liu, Dongqi, et al.
Published: (2025)
by: Liu, Dongqi, et al.
Published: (2025)
System-1.x: Learning to Balance Fast and Slow Planning with Language Models
by: Saha, Swarnadeep, et al.
Published: (2024)
by: Saha, Swarnadeep, et al.
Published: (2024)
Offline Learning and Forgetting for Reasoning with Large Language Models
by: Ni, Tianwei, et al.
Published: (2025)
by: Ni, Tianwei, et al.
Published: (2025)
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
by: Zhang, Yinger, et al.
Published: (2026)
by: Zhang, Yinger, et al.
Published: (2026)
Can Large Language Models Reason and Plan?
by: Kambhampati, Subbarao
Published: (2024)
by: Kambhampati, Subbarao
Published: (2024)
Iterative Deployment Improves Planning Skills in LLMs
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker Environments
by: Xia, Boyang, et al.
Published: (2026)
by: Xia, Boyang, et al.
Published: (2026)
Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
Offline Reinforcement Learning with Universal Horizon Models
by: Chung, Hojun, et al.
Published: (2026)
by: Chung, Hojun, et al.
Published: (2026)
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
by: Liu, Shih-Yang, et al.
Published: (2025)
by: Liu, Shih-Yang, et al.
Published: (2025)
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
by: Wang, Shenzhi, et al.
Published: (2025)
by: Wang, Shenzhi, et al.
Published: (2025)
Similar Items
-
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
by: Wang, Zehong, et al.
Published: (2026) -
LLoCO: Learning Long Contexts Offline
by: Tan, Sijun, et al.
Published: (2024) -
Guiding Language Model Reasoning with Planning Tokens
by: Wang, Xinyi, et al.
Published: (2023) -
Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
by: Ustaomeroglu, Muhammed, et al.
Published: (2025) -
On the Ability of Transformers to Verify Plans
by: Sarrof, Yash, et al.
Published: (2026)