GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Sijia, Huang, Yuchen, Liu, Zifan, Li, Yanping, Fu, Jingjing, Zhao, Li, Bian, Jiang, Zhang, Ling, Zhang, Jun, Wang, Rui
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914566931218432
author Li, Sijia
Huang, Yuchen
Liu, Zifan
Li, Yanping
Fu, Jingjing
Zhao, Li
Bian, Jiang
Zhang, Ling
Zhang, Jun
Wang, Rui
author_facet Li, Sijia
Huang, Yuchen
Liu, Zifan
Li, Yanping
Fu, Jingjing
Zhao, Li
Bian, Jiang
Zhang, Ling
Zhang, Jun
Wang, Rui
contents Reinforcement learning has become a widely used post-training approach for LLM agents, where training commonly relies on outcome-level rewards that provide only coarse supervision. While finer-grained credit assignment is promising for effective policy updates, obtaining reliable local credit and assigning it to the right parts of the long-horizon trajectory remains an open challenge. In this paper, we propose Granularity-adaptivE Advantage Reweighting (GEAR), an adaptive-granularity credit assignment framework that reshapes the trajectory-level GRPO advantage using token- and segment-level signals derived from self-distillation. GEAR compares an on-policy student with a ground-truth-conditioned teacher to obtain a reference-guided divergence signal for identifying adaptive segment boundaries and modulating local advantage weights. This divergence often spikes at the onset of a semantic deviation, while later tokens in the same autoregressive continuation may return to low divergence. GEAR therefore treats such spikes as anchors for adaptive credit regions: where the student remains aligned with the teacher, token-level resolution is preserved; where it departs, GEAR groups the corresponding continuation into an adaptive segment and uses the divergence at the departure point to modulate the segment' s advantage. Experiments across eight mathematical reasoning and agentic tool-use benchmarks with Qwen3 4B and 8B models show that GEAR consistently outperforms standard GRPO, self-distillation-only baselines, and token- or turn-level credit-assignment methods. The gains are especially strong on benchmarks with lower GRPO baseline accuracy, reaching up to around 20\% over GRPO, suggesting that the proposed adaptive reweighting scheme is especially useful in more challenging long-horizon settings.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11853
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
Li, Sijia
Huang, Yuchen
Liu, Zifan
Li, Yanping
Fu, Jingjing
Zhao, Li
Bian, Jiang
Zhang, Ling
Zhang, Jun
Wang, Rui
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement learning has become a widely used post-training approach for LLM agents, where training commonly relies on outcome-level rewards that provide only coarse supervision. While finer-grained credit assignment is promising for effective policy updates, obtaining reliable local credit and assigning it to the right parts of the long-horizon trajectory remains an open challenge. In this paper, we propose Granularity-adaptivE Advantage Reweighting (GEAR), an adaptive-granularity credit assignment framework that reshapes the trajectory-level GRPO advantage using token- and segment-level signals derived from self-distillation. GEAR compares an on-policy student with a ground-truth-conditioned teacher to obtain a reference-guided divergence signal for identifying adaptive segment boundaries and modulating local advantage weights. This divergence often spikes at the onset of a semantic deviation, while later tokens in the same autoregressive continuation may return to low divergence. GEAR therefore treats such spikes as anchors for adaptive credit regions: where the student remains aligned with the teacher, token-level resolution is preserved; where it departs, GEAR groups the corresponding continuation into an adaptive segment and uses the divergence at the departure point to modulate the segment' s advantage. Experiments across eight mathematical reasoning and agentic tool-use benchmarks with Qwen3 4B and 8B models show that GEAR consistently outperforms standard GRPO, self-distillation-only baselines, and token- or turn-level credit-assignment methods. The gains are especially strong on benchmarks with lower GRPO baseline accuracy, reaching up to around 20\% over GRPO, suggesting that the proposed adaptive reweighting scheme is especially useful in more challenging long-horizon settings.
title GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.11853