Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Siyu, Jiang, Yanbin, Sang, Hejian, Tang, Shao, Song, Qingquan, He, Biao, Jain, Rohit, Wang, Zhipeng, Geramifard, Alborz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
TIP: Token Importance in On-Policy Distillation
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Agentic Reinforcement Learning for Real-World Code Repair
von: Zhu, Siyu, et al.
Veröffentlicht: (2025)
von: Zhu, Siyu, et al.
Veröffentlicht: (2025)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
von: Chen, Xiwen, et al.
Veröffentlicht: (2026)
von: Chen, Xiwen, et al.
Veröffentlicht: (2026)
Sequential Decision-Making for Inline Text Autocomplete
von: Chitnis, Rohan, et al.
Veröffentlicht: (2024)
von: Chitnis, Rohan, et al.
Veröffentlicht: (2024)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
von: Chen, Wei-Rui, et al.
Veröffentlicht: (2025)
von: Chen, Wei-Rui, et al.
Veröffentlicht: (2025)
When should we prefer Decision Transformers for Offline Reinforcement Learning?
von: Bhargava, Prajjwal, et al.
Veröffentlicht: (2023)
von: Bhargava, Prajjwal, et al.
Veröffentlicht: (2023)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
von: Behdin, Kayhan, et al.
Veröffentlicht: (2025)
von: Behdin, Kayhan, et al.
Veröffentlicht: (2025)
AlphaPO: Reward Shape Matters for LLM Alignment
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Debunk the Myth of SFT Generalization
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2025)
SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning
von: Sikchi, Harshit, et al.
Veröffentlicht: (2023)
von: Sikchi, Harshit, et al.
Veröffentlicht: (2023)
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Not all tokens are needed(NAT): token efficient reinforcement learning
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
von: Duan, Xintong, et al.
Veröffentlicht: (2025)
von: Duan, Xintong, et al.
Veröffentlicht: (2025)
Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization
von: Su, Jianghao, et al.
Veröffentlicht: (2025)
von: Su, Jianghao, et al.
Veröffentlicht: (2025)
Aligning Diffusion Language Models via Unpaired Preference Optimization
von: Jindal, Vaibhav, et al.
Veröffentlicht: (2025)
von: Jindal, Vaibhav, et al.
Veröffentlicht: (2025)
Maximal subrings of certain non-commutative rings
von: Azarang, Alborz
Veröffentlicht: (2024)
von: Azarang, Alborz
Veröffentlicht: (2024)
Similar submodules of projective modules
von: Azarang, Alborz
Veröffentlicht: (2026)
von: Azarang, Alborz
Veröffentlicht: (2026)
The conductor ideals of maximal subrings in non-commutative rings
von: Azarang, Alborz
Veröffentlicht: (2024)
von: Azarang, Alborz
Veröffentlicht: (2024)
Classical ideals theory of maximal subrings in non-commutative rings
von: Azarang, Alborz
Veröffentlicht: (2024)
von: Azarang, Alborz
Veröffentlicht: (2024)
Non-commutative rings with infinitely many maximal subrings
von: Azarang, Alborz
Veröffentlicht: (2026)
von: Azarang, Alborz
Veröffentlicht: (2026)
Maximal subrings of division rings
von: Azarang, Alborz
Veröffentlicht: (2024)
von: Azarang, Alborz
Veröffentlicht: (2024)
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
von: Lucas, Ryan, et al.
Veröffentlicht: (2025)
von: Lucas, Ryan, et al.
Veröffentlicht: (2025)
Bayesian Preference Learning for Test-Time Steerable Reward Models
von: Hong, Jiwoo, et al.
Veröffentlicht: (2026)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2026)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2026)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2026)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
von: Nazir, Mohammad Saif, et al.
Veröffentlicht: (2025)
von: Nazir, Mohammad Saif, et al.
Veröffentlicht: (2025)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning
von: Dong, Xuanzhao, et al.
Veröffentlicht: (2026)
von: Dong, Xuanzhao, et al.
Veröffentlicht: (2026)
GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs
von: Feng, Tao, et al.
Veröffentlicht: (2026)
von: Feng, Tao, et al.
Veröffentlicht: (2026)
A Tactile-based Interactive Motion Planner for Robots in Unknown Cluttered Environments
von: Wang, Chengjin, et al.
Veröffentlicht: (2025)
von: Wang, Chengjin, et al.
Veröffentlicht: (2025)
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
von: Garg, Spandan, et al.
Veröffentlicht: (2026)
von: Garg, Spandan, et al.
Veröffentlicht: (2026)
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
von: Feng, Xiao, et al.
Veröffentlicht: (2026)
von: Feng, Xiao, et al.
Veröffentlicht: (2026)
Liger Kernel: Efficient Triton Kernels for LLM Training
von: Hsu, Pin-Lun, et al.
Veröffentlicht: (2024)
von: Hsu, Pin-Lun, et al.
Veröffentlicht: (2024)
Search-P1: Path-Centric Reward Shaping for Stable and Efficient Agentic RAG Training
von: Xia, Tianle, et al.
Veröffentlicht: (2026)
von: Xia, Tianle, et al.
Veröffentlicht: (2026)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
von: Groeneveld, Jan Niklas, et al.
Veröffentlicht: (2025)
von: Groeneveld, Jan Niklas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026) -
TIP: Token Importance in On-Policy Distillation
von: Xu, Yuanda, et al.
Veröffentlicht: (2026) -
Agentic Reinforcement Learning for Real-World Code Repair
von: Zhu, Siyu, et al.
Veröffentlicht: (2025) -
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
von: Chen, Xiwen, et al.
Veröffentlicht: (2026) -
Sequential Decision-Making for Inline Text Autocomplete
von: Chitnis, Rohan, et al.
Veröffentlicht: (2024)