The Path Not Taken: RLVR Provably Learns Off the Principals

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Hanqing, Zhang, Zhenyu, Huang, Hanxian, Su, DiJia, Liu, Zechun, Zhao, Jiawei, Fedorov, Igor, Pirsiavash, Hamed, Sha, Zhizhou, Lee, Jinwon, Pan, David Z., Wang, Zhangyang, Tian, Yuandong, Tai, Kai Sheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914151485407232
author Zhu, Hanqing
Zhang, Zhenyu
Huang, Hanxian
Su, DiJia
Liu, Zechun
Zhao, Jiawei
Fedorov, Igor
Pirsiavash, Hamed
Sha, Zhizhou
Lee, Jinwon
Pan, David Z.
Wang, Zhangyang
Tian, Yuandong
Tai, Kai Sheng
author_facet Zhu, Hanqing
Zhang, Zhenyu
Huang, Hanxian
Su, DiJia
Liu, Zechun
Zhao, Jiawei
Fedorov, Igor
Pirsiavash, Hamed
Sha, Zhizhou
Lee, Jinwon
Pan, David Z.
Wang, Zhangyang
Tian, Yuandong
Tai, Kai Sheng
contents Reinforcement Learning with Verifiable Rewards (RLVR) reliably improves the reasoning performance of large language models, yet it appears to modify only a small fraction of parameters. We revisit this paradox and show that sparsity is a surface artifact of a model-conditioned optimization bias: for a fixed pretrained model, updates consistently localize to preferred parameter regions, highly consistent across runs and largely invariant to datasets and RL recipes. We mechanistically explain these dynamics with a Three-Gate Theory: Gate I (KL Anchor) imposes a KL-constrained update; Gate II (Model Geometry) steers the step off principal directions into low-curvature, spectrum-preserving subspaces; and Gate III (Precision) hides micro-updates in non-preferred regions, making the off-principal bias appear as sparsity. We then validate this theory and, for the first time, provide a parameter-level characterization of RLVR's learning dynamics: RLVR learns off principal directions in weight space, achieving gains via minimal spectral drift, reduced principal-subspace rotation, and off-principal update alignment. In contrast, SFT targets principal weights, distorts the spectrum, and even lags RLVR. Together, these results provide the first parameter-space account of RLVR's training dynamics, revealing clear regularities in how parameters evolve. Crucially, we show that RL operates in a distinct optimization regime from SFT, so directly adapting SFT-era parameter-efficient fine-tuning (PEFT) methods can be flawed, as evidenced by our case studies on advanced sparse fine-tuning and LoRA variants. We hope this work charts a path toward a white-box understanding of RLVR and the design of geometry-aware, RLVR-native learning algorithms, rather than repurposed SFT-era heuristics.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08567
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Path Not Taken: RLVR Provably Learns Off the Principals
Zhu, Hanqing
Zhang, Zhenyu
Huang, Hanxian
Su, DiJia
Liu, Zechun
Zhao, Jiawei
Fedorov, Igor
Pirsiavash, Hamed
Sha, Zhizhou
Lee, Jinwon
Pan, David Z.
Wang, Zhangyang
Tian, Yuandong
Tai, Kai Sheng
Machine Learning
Artificial Intelligence
Reinforcement Learning with Verifiable Rewards (RLVR) reliably improves the reasoning performance of large language models, yet it appears to modify only a small fraction of parameters. We revisit this paradox and show that sparsity is a surface artifact of a model-conditioned optimization bias: for a fixed pretrained model, updates consistently localize to preferred parameter regions, highly consistent across runs and largely invariant to datasets and RL recipes. We mechanistically explain these dynamics with a Three-Gate Theory: Gate I (KL Anchor) imposes a KL-constrained update; Gate II (Model Geometry) steers the step off principal directions into low-curvature, spectrum-preserving subspaces; and Gate III (Precision) hides micro-updates in non-preferred regions, making the off-principal bias appear as sparsity. We then validate this theory and, for the first time, provide a parameter-level characterization of RLVR's learning dynamics: RLVR learns off principal directions in weight space, achieving gains via minimal spectral drift, reduced principal-subspace rotation, and off-principal update alignment. In contrast, SFT targets principal weights, distorts the spectrum, and even lags RLVR. Together, these results provide the first parameter-space account of RLVR's training dynamics, revealing clear regularities in how parameters evolve. Crucially, we show that RL operates in a distinct optimization regime from SFT, so directly adapting SFT-era parameter-efficient fine-tuning (PEFT) methods can be flawed, as evidenced by our case studies on advanced sparse fine-tuning and LoRA variants. We hope this work charts a path toward a white-box understanding of RLVR and the design of geometry-aware, RLVR-native learning algorithms, rather than repurposed SFT-era heuristics.
title The Path Not Taken: RLVR Provably Learns Off the Principals
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.08567