Self-Distilled RLVR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Chenxu, Qin, Chuanyu, Si, Qingyi, Chen, Minghui, Gu, Naibin, Yao, Dingyu, Lin, Zheng, Wang, Weiping, Wang, Jiaqi, Duan, Nan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Co-Evolving Policy Distillation
von: Gu, Naibin, et al.
Veröffentlicht: (2026)
von: Gu, Naibin, et al.
Veröffentlicht: (2026)
Near-Future Policy Optimization
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026)
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026)
EasyVideoR1: Easier RL for Video Understanding
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026)
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026)
System 1&2 Synergy via Dynamic Model Interpolation
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
Test-time Prompt Intervention
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
Orthogonal Finetuning for Direct Preference Optimization
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
von: Yang, Chenxu, et al.
Veröffentlicht: (2024)
Are Large Language Models Table-based Fact-Checkers?
von: Zhang, Hanwen, et al.
Veröffentlicht: (2024)
von: Zhang, Hanwen, et al.
Veröffentlicht: (2024)
Beyond the Covariance Trap: Unlocking Generalization in Same-Subject Knowledge Editing for Large Language Models
von: Liu, Xiyu, et al.
Veröffentlicht: (2026)
von: Liu, Xiyu, et al.
Veröffentlicht: (2026)
Online Self-Calibration Against Hallucination in Vision-Language Models
von: Chen, Minghui, et al.
Veröffentlicht: (2026)
von: Chen, Minghui, et al.
Veröffentlicht: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
Dynamic Early Exit in Reasoning Models
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
The Unlearnability Phenomenon in RLVR for Language Models
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?
von: Wang, Yatong, et al.
Veröffentlicht: (2026)
von: Wang, Yatong, et al.
Veröffentlicht: (2026)
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
Linear Dynamics in the RLVR Training of Large Language Models
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
A Multi-Task Role-Playing Agent Capable of Imitating Character Linguistic Styles
von: Chen, Siyuan, et al.
Veröffentlicht: (2024)
von: Chen, Siyuan, et al.
Veröffentlicht: (2024)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
A Closer Look into LLMs for Table Understanding
von: Wang, Jia, et al.
Veröffentlicht: (2026)
von: Wang, Jia, et al.
Veröffentlicht: (2026)
Causal Path Alignment: Anchoring the Optimization Trajectory for Controllable In-Parameter Knowledge Editing
von: Liu, Xiyu, et al.
Veröffentlicht: (2025)
von: Liu, Xiyu, et al.
Veröffentlicht: (2025)
Sparse Attention across Multiple-context KV Cache
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
Self-Distilled Agentic Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
Weights-Rotated Preference Optimization for Large Language Models
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
Efficient RLVR Training via Weighted Mutual Information Data Selection
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
Multilingual Safety Alignment via Self-Distillation
von: Qin, Ruiyang, et al.
Veröffentlicht: (2026)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2026)
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
von: Wei, Zhepei, et al.
Veröffentlicht: (2026)
von: Wei, Zhepei, et al.
Veröffentlicht: (2026)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
von: Zhai, Zepeng, et al.
Veröffentlicht: (2026)
von: Zhai, Zepeng, et al.
Veröffentlicht: (2026)
TableDreamer: Progressive and Weakness-guided Data Synthesis from Scratch for Table Instruction Tuning
von: Zheng, Mingyu, et al.
Veröffentlicht: (2025)
von: Zheng, Mingyu, et al.
Veröffentlicht: (2025)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Co-Evolving Policy Distillation
von: Gu, Naibin, et al.
Veröffentlicht: (2026) -
Near-Future Policy Optimization
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026) -
EasyVideoR1: Easier RL for Video Understanding
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026) -
System 1&2 Synergy via Dynamic Model Interpolation
von: Yang, Chenxu, et al.
Veröffentlicht: (2026) -
Test-time Prompt Intervention
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)