TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiaqi, Zhang, Wenhao, Shi, Weijie, Li, Yaliang, Cheng, James |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
by: Luo, Haohao, et al.
Published: (2026)
by: Luo, Haohao, et al.
Published: (2026)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
by: Gao, Zhaolin, et al.
Published: (2024)
by: Gao, Zhaolin, et al.
Published: (2024)
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
by: Cai, Yuzhu, et al.
Published: (2026)
by: Cai, Yuzhu, et al.
Published: (2026)
Agent-Oriented Planning in Multi-Agent Systems
by: Li, Ao, et al.
Published: (2024)
by: Li, Ao, et al.
Published: (2024)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
by: Ma, Chang, et al.
Published: (2024)
by: Ma, Chang, et al.
Published: (2024)
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
by: Wang, Shumin, et al.
Published: (2026)
by: Wang, Shumin, et al.
Published: (2026)
R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification
by: Shi, Weijie, et al.
Published: (2026)
by: Shi, Weijie, et al.
Published: (2026)
Temp-R1: A Unified Autonomous Agent for Complex Temporal KGQA via Reverse Curriculum Reinforcement Learning
by: Gong, Zhaoyan, et al.
Published: (2026)
by: Gong, Zhaoyan, et al.
Published: (2026)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
by: Wang, Jianze, et al.
Published: (2026)
by: Wang, Jianze, et al.
Published: (2026)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
by: Yao, Chaorui, et al.
Published: (2025)
by: Yao, Chaorui, et al.
Published: (2025)
Reinforcement Learning with Curriculum-inspired Adaptive Direct Policy Guidance for Truck Dispatching
by: Meng, Shi, et al.
Published: (2025)
by: Meng, Shi, et al.
Published: (2025)
Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
by: Lauffer, Niklas, et al.
Published: (2025)
by: Lauffer, Niklas, et al.
Published: (2025)
Collaborative Adaptive Curriculum for Progressive Knowledge Distillation
by: Liu, Jing, et al.
Published: (2026)
by: Liu, Jing, et al.
Published: (2026)
MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
by: Duan, Wenchang, et al.
Published: (2026)
by: Duan, Wenchang, et al.
Published: (2026)
Learning Progress Driven Multi-Agent Curriculum
by: Zhao, Wenshuai, et al.
Published: (2022)
by: Zhao, Wenshuai, et al.
Published: (2022)
Rethinking Spatio-Temporal Transformer for Traffic Prediction:Multi-level Multi-view Augmented Learning Framework
by: Lin, Jiaqi, et al.
Published: (2024)
by: Lin, Jiaqi, et al.
Published: (2024)
OM2P: Offline Multi-Agent Mean-Flow Policy
by: Li, Zhuoran, et al.
Published: (2025)
by: Li, Zhuoran, et al.
Published: (2025)
Extreme Region Policy Distillation
by: Chen, Changyu, et al.
Published: (2026)
by: Chen, Changyu, et al.
Published: (2026)
PTCL: Pseudo-Label Temporal Curriculum Learning for Label-Limited Dynamic Graph
by: Zhang, Shengtao, et al.
Published: (2025)
by: Zhang, Shengtao, et al.
Published: (2025)
Bidirectional Distillation: A Mixed-Play Framework for Multi-Agent Generalizable Behaviors
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
by: Yu, Bowen, et al.
Published: (2026)
by: Yu, Bowen, et al.
Published: (2026)
Curriculum Negative Mining For Temporal Networks
by: Chen, Ziyue, et al.
Published: (2024)
by: Chen, Ziyue, et al.
Published: (2024)
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
by: Cheng, Gang, et al.
Published: (2025)
by: Cheng, Gang, et al.
Published: (2025)
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
MTDrive: Multi-turn Interactive Reinforcement Learning for Autonomous Driving
by: Li, Xidong, et al.
Published: (2026)
by: Li, Xidong, et al.
Published: (2026)
Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
Proximal Policy Distillation
by: Spigler, Giacomo
Published: (2024)
by: Spigler, Giacomo
Published: (2024)
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Interpretable Hybrid-Rule Temporal Point Processes
by: Cao, Yunyang, et al.
Published: (2025)
by: Cao, Yunyang, et al.
Published: (2025)
Review of Data-centric Time Series Analysis from Sample, Feature, and Period
by: Sun, Chenxi, et al.
Published: (2024)
by: Sun, Chenxi, et al.
Published: (2024)
TIP: Token Importance in On-Policy Distillation
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Curriculum Learning-Guided Progressive Distillation in Large Language Models
by: Cao, Jincheng, et al.
Published: (2026)
by: Cao, Jincheng, et al.
Published: (2026)
Similar Items
-
IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
by: Luo, Haohao, et al.
Published: (2026) -
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
by: Zhang, Wenhao, et al.
Published: (2025) -
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
by: Wang, Hao, et al.
Published: (2026) -
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
by: Li, Yu, et al.
Published: (2026) -
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
by: Gao, Zhaolin, et al.
Published: (2024)