Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Dou, Zhihao, Zhao, Qinjian, Wan, Zhongwei, Zhang, Dinggen, Wang, Weida, Raiyan, Towsif, Chen, Benteng, Pan, Qingtao, Ouyang, Yang, Song, Chaoda, Gao, Zhiqiang, Zhang, Shufei, Biswas, Sumon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoRe-Code: Collaborative Reinforcement Learning for Code Generation
by: Dou, Zhihao, et al.
Published: (2026)
by: Dou, Zhihao, et al.
Published: (2026)
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
by: Chen, Benteng, et al.
Published: (2026)
by: Chen, Benteng, et al.
Published: (2026)
DSADF: Thinking Fast and Slow for Decision Making
by: Dou, Zhihao, et al.
Published: (2025)
by: Dou, Zhihao, et al.
Published: (2025)
SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
by: Zhao, Qinjian, et al.
Published: (2025)
by: Zhao, Qinjian, et al.
Published: (2025)
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
by: Wan, Zhongwei, et al.
Published: (2025)
by: Wan, Zhongwei, et al.
Published: (2025)
Frequency-Modulated Visual Restoration for Matryoshka Large Multimodal Models
by: Pan, Qingtao, et al.
Published: (2026)
by: Pan, Qingtao, et al.
Published: (2026)
Escaping Optimization Stagnation: Taking Steps Beyond Task Arithmetic via Difference Vectors
by: Wang, Jinping, et al.
Published: (2025)
by: Wang, Jinping, et al.
Published: (2025)
Surgical Action Planning with Large Language Models
by: Xu, Mengya, et al.
Published: (2025)
by: Xu, Mengya, et al.
Published: (2025)
Chem-R: Learning to Reason as a Chemist
by: Wang, Weida, et al.
Published: (2025)
by: Wang, Weida, et al.
Published: (2025)
Equivariant Action Sampling for Reinforcement Learning and Planning
by: Zhao, Linfeng, et al.
Published: (2024)
by: Zhao, Linfeng, et al.
Published: (2024)
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
by: Nazarenus, Eric, et al.
Published: (2026)
by: Nazarenus, Eric, et al.
Published: (2026)
The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models
by: Hasan, Alif Al, et al.
Published: (2026)
by: Hasan, Alif Al, et al.
Published: (2026)
What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
by: Hasan, Alif Al, et al.
Published: (2026)
by: Hasan, Alif Al, et al.
Published: (2026)
Reinforced Reasoning for Embodied Planning
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
by: Li, Jiatong, et al.
Published: (2025)
by: Li, Jiatong, et al.
Published: (2025)
Efficient Soft Actor-Critic with LLM-Based Action-Level Guidance for Continuous Control
by: Ma, Hao, et al.
Published: (2026)
by: Ma, Hao, et al.
Published: (2026)
Enhancing Numerical Reasoning with the Guidance of Reliable Reasoning Processes
by: Wang, Dingzirui, et al.
Published: (2024)
by: Wang, Dingzirui, et al.
Published: (2024)
PolyReal: A Benchmark for Real-World Polymer Science Workflows
by: Liu, Wanhao, et al.
Published: (2026)
by: Liu, Wanhao, et al.
Published: (2026)
Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models
by: Raiyan, Syed Rifat
Published: (2026)
by: Raiyan, Syed Rifat
Published: (2026)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
ACPBench: Reasoning about Action, Change, and Planning
by: Kokel, Harsha, et al.
Published: (2024)
by: Kokel, Harsha, et al.
Published: (2024)
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
by: Xu, Mengya, et al.
Published: (2025)
by: Xu, Mengya, et al.
Published: (2025)
Plan-MCTS: Plan Exploration for Action Exploitation in Web Navigation
by: Zhang, Weiming, et al.
Published: (2026)
by: Zhang, Weiming, et al.
Published: (2026)
AOT*: Efficient Synthesis Planning via LLM-Empowered AND-OR Tree Search
by: Song, Xiaozhuang, et al.
Published: (2025)
by: Song, Xiaozhuang, et al.
Published: (2025)
SAMG: Offline-to-Online Reinforcement Learning via State-Action-Conditional Offline Model Guidance
by: Zhang, Liyu, et al.
Published: (2024)
by: Zhang, Liyu, et al.
Published: (2024)
QCBench: Evaluating Large Language Models on Domain-Specific Quantitative Chemistry
by: Xie, Jiaqing, et al.
Published: (2025)
by: Xie, Jiaqing, et al.
Published: (2025)
DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning
by: Wan, Zhongwei, et al.
Published: (2026)
by: Wan, Zhongwei, et al.
Published: (2026)
Planning-Augmented Sampling with Early Guidance for High-Reward Discovery
by: Zhu, Rui, et al.
Published: (2025)
by: Zhu, Rui, et al.
Published: (2025)
Reinforced Reasoning for End-to-End Retrosynthetic Planning
by: Zuo, Chenyang, et al.
Published: (2026)
by: Zuo, Chenyang, et al.
Published: (2026)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
ACPBench Hard: Unrestrained Reasoning about Action, Change, and Planning
by: Kokel, Harsha, et al.
Published: (2025)
by: Kokel, Harsha, et al.
Published: (2025)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
by: Huang, Haoyu, et al.
Published: (2026)
by: Huang, Haoyu, et al.
Published: (2026)
Beyond Distributions: Geometric Action Control for Continuous Reinforcement Learning
by: Lin, Zhihao
Published: (2025)
by: Lin, Zhihao
Published: (2025)
On Reasoning Strength Planning in Large Reasoning Models
by: Sheng, Leheng, et al.
Published: (2025)
by: Sheng, Leheng, et al.
Published: (2025)
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
by: Li, Junxian, et al.
Published: (2026)
by: Li, Junxian, et al.
Published: (2026)
Path Database Guidance for Motion Planning
by: Attali, Amnon, et al.
Published: (2025)
by: Attali, Amnon, et al.
Published: (2025)
Reinforced Context Order Recovery for Adaptive Reasoning and Planning
by: Ma, Long, et al.
Published: (2025)
by: Ma, Long, et al.
Published: (2025)
Scheduling Drone and Mobile Charger via Hybrid-Action Deep Reinforcement Learning
by: Dou, Jizhe, et al.
Published: (2024)
by: Dou, Jizhe, et al.
Published: (2024)
Strategy Executability in Mathematical Reasoning: Leveraging Human-Model Differences for Effective Guidance
by: Liang, Weida, et al.
Published: (2026)
by: Liang, Weida, et al.
Published: (2026)
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
by: Yang, Wang, et al.
Published: (2026)
by: Yang, Wang, et al.
Published: (2026)
Similar Items
-
CoRe-Code: Collaborative Reinforcement Learning for Code Generation
by: Dou, Zhihao, et al.
Published: (2026) -
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
by: Chen, Benteng, et al.
Published: (2026) -
DSADF: Thinking Fast and Slow for Decision Making
by: Dou, Zhihao, et al.
Published: (2025) -
SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
by: Zhao, Qinjian, et al.
Published: (2025) -
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
by: Wan, Zhongwei, et al.
Published: (2025)