Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yu, Tang, Sizhe, Lan, Tian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
von: Zhang, Zuyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Zuyuan, et al.
Veröffentlicht: (2026)
Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics
von: Zhang, Zuyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Zuyuan, et al.
Veröffentlicht: (2026)
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
Annealing Self-Distillation Rectification Improves Adversarial Training
von: Wu, Yu-Yu, et al.
Veröffentlicht: (2023)
von: Wu, Yu-Yu, et al.
Veröffentlicht: (2023)
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
von: Liu, Zongkai, et al.
Veröffentlicht: (2024)
von: Liu, Zongkai, et al.
Veröffentlicht: (2024)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
von: Ma, Chang, et al.
Veröffentlicht: (2024)
von: Ma, Chang, et al.
Veröffentlicht: (2024)
Geometric Manifold Rectification for Imbalanced Learning
von: Wang, Xubin, et al.
Veröffentlicht: (2026)
von: Wang, Xubin, et al.
Veröffentlicht: (2026)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
Segment-Aligned Policy Optimization for Multi-Modal Reasoning
von: Gao, Lei, et al.
Veröffentlicht: (2026)
von: Gao, Lei, et al.
Veröffentlicht: (2026)
Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
von: Lauffer, Niklas, et al.
Veröffentlicht: (2025)
von: Lauffer, Niklas, et al.
Veröffentlicht: (2025)
FedSDR: Federated Self-Distillation with Rectification
von: Ren, Ziheng, et al.
Veröffentlicht: (2026)
von: Ren, Ziheng, et al.
Veröffentlicht: (2026)
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
von: Zhou, Huilin, et al.
Veröffentlicht: (2026)
von: Zhou, Huilin, et al.
Veröffentlicht: (2026)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Scaling Graph Chain-of-Thought Reasoning: A Multi-Agent Framework with Efficient LLM Serving
von: Huan, Chengying, et al.
Veröffentlicht: (2025)
von: Huan, Chengying, et al.
Veröffentlicht: (2025)
Efficient Rectification of Neuro-Symbolic Reasoning Inconsistencies by Abductive Reflection
von: Hu, Wen-Chao, et al.
Veröffentlicht: (2024)
von: Hu, Wen-Chao, et al.
Veröffentlicht: (2024)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
von: Li, Xuan, et al.
Veröffentlicht: (2026)
von: Li, Xuan, et al.
Veröffentlicht: (2026)
Federation over Text: Insight Sharing for Multi-Agent Reasoning
von: Yao, Dixi, et al.
Veröffentlicht: (2026)
von: Yao, Dixi, et al.
Veröffentlicht: (2026)
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
von: Feng, Xinshun, et al.
Veröffentlicht: (2026)
von: Feng, Xinshun, et al.
Veröffentlicht: (2026)
Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use Agents
von: Tang, Sizhe, et al.
Veröffentlicht: (2026)
von: Tang, Sizhe, et al.
Veröffentlicht: (2026)
Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy Churn
von: Tang, Hongyao, et al.
Veröffentlicht: (2024)
von: Tang, Hongyao, et al.
Veröffentlicht: (2024)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2026)
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2026)
MALinZero: Efficient Low-Dimensional Search for Mastering Complex Multi-Agent Planning
von: Tang, Sizhe, et al.
Veröffentlicht: (2025)
von: Tang, Sizhe, et al.
Veröffentlicht: (2025)
A Rectification-Based Approach for Distilling Boosted Trees into Decision Trees
von: Audemard, Gilles, et al.
Veröffentlicht: (2025)
von: Audemard, Gilles, et al.
Veröffentlicht: (2025)
PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind
von: Yu, Yajie, et al.
Veröffentlicht: (2025)
von: Yu, Yajie, et al.
Veröffentlicht: (2025)
TreeRPO: Tree Relative Policy Optimization
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
Multi-Agent Reinforcement Learning for Unmanned Aerial Vehicle Coordination by Multi-Critic Policy Gradient Optimization
von: Alon, Yoav, et al.
Veröffentlicht: (2020)
von: Alon, Yoav, et al.
Veröffentlicht: (2020)
ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical Dialogue
von: Cao, Ruike, et al.
Veröffentlicht: (2026)
von: Cao, Ruike, et al.
Veröffentlicht: (2026)
LEPO: Latent Reasoning Policy Optimization for Large Language Models
von: Zhou, Yuyan, et al.
Veröffentlicht: (2026)
von: Zhou, Yuyan, et al.
Veröffentlicht: (2026)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
SPAR: Support-Preserving Action Rectification
von: Zhao, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhao, Jiaxin, et al.
Veröffentlicht: (2026)
FORLER: Federated Offline Reinforcement Learning with Q-Ensemble and Actor Rectification
von: Qiao, Nan, et al.
Veröffentlicht: (2026)
von: Qiao, Nan, et al.
Veröffentlicht: (2026)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
von: Reddi, Aryaman, et al.
Veröffentlicht: (2025)
von: Reddi, Aryaman, et al.
Veröffentlicht: (2025)
Policy Guided Tree Search for Enhanced LLM Reasoning
von: Li, Yang
Veröffentlicht: (2025)
von: Li, Yang
Veröffentlicht: (2025)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
von: Zhang, Zuyuan, et al.
Veröffentlicht: (2026) -
Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics
von: Zhang, Zuyuan, et al.
Veröffentlicht: (2026) -
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026) -
InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
von: Li, Yu, et al.
Veröffentlicht: (2025) -
Annealing Self-Distillation Rectification Improves Adversarial Training
von: Wu, Yu-Yu, et al.
Veröffentlicht: (2023)