CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Shijie, Sun, Guohao, Zhang, Kevin, Guo, Xiang, Guo, Rujun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Driven Policy Diffusion: Enhancing Generalization in Offline Reinforcement Learning
by: Zhang, Hanping, et al.
Published: (2025)
by: Zhang, Hanping, et al.
Published: (2025)
Self-Evolving Curriculum for LLM Reasoning
by: Chen, Xiaoyin, et al.
Published: (2025)
by: Chen, Xiaoyin, et al.
Published: (2025)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation
by: Zhang, Enci, et al.
Published: (2025)
by: Zhang, Enci, et al.
Published: (2025)
ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
by: Wang, Yunhao, et al.
Published: (2025)
by: Wang, Yunhao, et al.
Published: (2025)
SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning
by: Li, Chenyi, et al.
Published: (2026)
by: Li, Chenyi, et al.
Published: (2026)
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
by: Wu, Xian, et al.
Published: (2026)
by: Wu, Xian, et al.
Published: (2026)
Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning
by: Yin, Shouyu, et al.
Published: (2026)
by: Yin, Shouyu, et al.
Published: (2026)
Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning
by: Sun, Yiliu, et al.
Published: (2025)
by: Sun, Yiliu, et al.
Published: (2025)
ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization
by: Zhang, Shijie, et al.
Published: (2026)
by: Zhang, Shijie, et al.
Published: (2026)
MemPO: Self-Memory Policy Optimization for Long-Horizon Agents
by: Li, Ruoran, et al.
Published: (2026)
by: Li, Ruoran, et al.
Published: (2026)
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning
by: Li, Xuchen, et al.
Published: (2026)
by: Li, Xuchen, et al.
Published: (2026)
Code Evolution for Control: Synthesizing Policies via LLM-Driven Evolutionary Search
by: Guo, Ping, et al.
Published: (2026)
by: Guo, Ping, et al.
Published: (2026)
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2024)
by: Xi, Zhiheng, et al.
Published: (2024)
Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning
by: Zhang, Shijie, et al.
Published: (2026)
by: Zhang, Shijie, et al.
Published: (2026)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
Policy of Thoughts: Scaling LLM Reasoning via Test-time Policy Evolution
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
by: Parashar, Shubham, et al.
Published: (2025)
by: Parashar, Shubham, et al.
Published: (2025)
Adversarial Curriculum Graph Contrastive Learning with Pair-wise Augmentation
by: Zhao, Xinjian, et al.
Published: (2024)
by: Zhao, Xinjian, et al.
Published: (2024)
Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
by: Wang, Ziyan, et al.
Published: (2025)
by: Wang, Ziyan, et al.
Published: (2025)
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
by: Ke, Junlong, et al.
Published: (2026)
by: Ke, Junlong, et al.
Published: (2026)
2D-Curri-DPO: Two-Dimensional Curriculum Learning for Direct Preference Optimization
by: Li, Mengyang, et al.
Published: (2025)
by: Li, Mengyang, et al.
Published: (2025)
DARC: Decoupled Asymmetric Reasoning Curriculum for LLM Evolution
by: Fan, Shengda, et al.
Published: (2026)
by: Fan, Shengda, et al.
Published: (2026)
COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
by: Wan, Guangya, et al.
Published: (2025)
by: Wan, Guangya, et al.
Published: (2025)
Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
by: Guo, Ruocheng, et al.
Published: (2026)
by: Guo, Ruocheng, et al.
Published: (2026)
Analyzing and Internalizing Complex Policy Documents for LLM Agents
by: Liu, Jiateng, et al.
Published: (2025)
by: Liu, Jiateng, et al.
Published: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
Reinforcement Learning with Curriculum-inspired Adaptive Direct Policy Guidance for Truck Dispatching
by: Meng, Shi, et al.
Published: (2025)
by: Meng, Shi, et al.
Published: (2025)
Calibration-Aware Policy Optimization for Reasoning LLMs
by: Wang, Ziqi, et al.
Published: (2026)
by: Wang, Ziqi, et al.
Published: (2026)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool Learning
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
Multi-objective Optimization by Learning Space Partitions
by: Zhao, Yiyang, et al.
Published: (2021)
by: Zhao, Yiyang, et al.
Published: (2021)
Offline Policy Optimization with Posterior Sampling
by: Lin, Hongqiang, et al.
Published: (2026)
by: Lin, Hongqiang, et al.
Published: (2026)
CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs
by: Zeng, Yongcheng, et al.
Published: (2025)
by: Zeng, Yongcheng, et al.
Published: (2025)
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
by: Hong, Minjie, et al.
Published: (2025)
by: Hong, Minjie, et al.
Published: (2025)
Segment-Aligned Policy Optimization for Multi-Modal Reasoning
by: Gao, Lei, et al.
Published: (2026)
by: Gao, Lei, et al.
Published: (2026)
Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning
by: Rui, Shaohao, et al.
Published: (2025)
by: Rui, Shaohao, et al.
Published: (2025)
ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning
by: Gao, Jingyue, et al.
Published: (2026)
by: Gao, Jingyue, et al.
Published: (2026)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy Optimization
by: Zhang, Siwei, et al.
Published: (2026)
by: Zhang, Siwei, et al.
Published: (2026)
Similar Items
-
LLM-Driven Policy Diffusion: Enhancing Generalization in Offline Reinforcement Learning
by: Zhang, Hanping, et al.
Published: (2025) -
Self-Evolving Curriculum for LLM Reasoning
by: Chen, Xiaoyin, et al.
Published: (2025) -
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026) -
Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation
by: Zhang, Enci, et al.
Published: (2025) -
ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
by: Wang, Yunhao, et al.
Published: (2025)