From Correction to Mastery: Reinforced Distillation of Large Language Model Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lyu, Yuanjie, Wang, Chengyu, Huang, Jun, Xu, Tong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916997713887232
author Lyu, Yuanjie
Wang, Chengyu
Huang, Jun
Xu, Tong
author_facet Lyu, Yuanjie
Wang, Chengyu
Huang, Jun
Xu, Tong
contents Large Language Model agents excel at solving complex tasks through iterative reasoning and tool use, but typically depend on ultra-large, costly backbones. Existing distillation approaches train smaller students to imitate full teacher trajectories, yet reasoning and knowledge gaps between the teacher and student can cause compounding errors. We propose SCoRe, a student-centered framework in which the student generates training trajectories and the teacher corrects only the earliest error, producing training data matched to the student's ability and exposing specific weaknesses. The student is first fine-tuned on corrected trajectories. Subsequently, short-horizon reinforcement learning starts from the verified prefix preceding the earliest error, with target rewards assigned at that step. This design encourages autonomous problem-solving beyond imitation and enhances training stability. On 12 challenging benchmarks, a 7B-parameter student distilled with SCoRe matches the agentic performance of a 72B-parameter teacher.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14257
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Correction to Mastery: Reinforced Distillation of Large Language Model Agents
Lyu, Yuanjie
Wang, Chengyu
Huang, Jun
Xu, Tong
Computation and Language
Artificial Intelligence
Large Language Model agents excel at solving complex tasks through iterative reasoning and tool use, but typically depend on ultra-large, costly backbones. Existing distillation approaches train smaller students to imitate full teacher trajectories, yet reasoning and knowledge gaps between the teacher and student can cause compounding errors. We propose SCoRe, a student-centered framework in which the student generates training trajectories and the teacher corrects only the earliest error, producing training data matched to the student's ability and exposing specific weaknesses. The student is first fine-tuned on corrected trajectories. Subsequently, short-horizon reinforcement learning starts from the verified prefix preceding the earliest error, with target rewards assigned at that step. This design encourages autonomous problem-solving beyond imitation and enhances training stability. On 12 challenging benchmarks, a 7B-parameter student distilled with SCoRe matches the agentic performance of a 72B-parameter teacher.
title From Correction to Mastery: Reinforced Distillation of Large Language Model Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.14257