Evaluating Parameter Efficient Methods for RLVR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Qingyu, Wu, Yulun, Shen, Zhennan, Li, Sunbowen, Wang, Zhilin, Li, Yanshu, Leong, Chak Tou, Kang, Jiale, Gu, Jinjin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering
von: Leong, Chak Tou, et al.
Veröffentlicht: (2026)
von: Leong, Chak Tou, et al.
Veröffentlicht: (2026)
Data-Efficient RLVR via Off-Policy Influence Guidance
von: Zhu, Erle, et al.
Veröffentlicht: (2025)
von: Zhu, Erle, et al.
Veröffentlicht: (2025)
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
von: Leong, Chak Tou, et al.
Veröffentlicht: (2025)
von: Leong, Chak Tou, et al.
Veröffentlicht: (2025)
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
Probing the Difficulty Perception Mechanism of Large Language Models
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
Protein as a Second Language for LLMs
von: Chen, Xinhui, et al.
Veröffentlicht: (2025)
von: Chen, Xinhui, et al.
Veröffentlicht: (2025)
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
Self-Distilled RLVR
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
von: Wang, Shuxun, et al.
Veröffentlicht: (2025)
von: Wang, Shuxun, et al.
Veröffentlicht: (2025)
Not only where, But when: Temporal Scheduling for RLVR
von: Zhang, Jinghao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinghao, et al.
Veröffentlicht: (2026)
RLVR-World: Training World Models with Reinforcement Learning
von: Wu, Jialong, et al.
Veröffentlicht: (2025)
von: Wu, Jialong, et al.
Veröffentlicht: (2025)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
von: Li, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Li, Kaiyuan, et al.
Veröffentlicht: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
von: Miao, Yuchun, et al.
Veröffentlicht: (2026)
von: Miao, Yuchun, et al.
Veröffentlicht: (2026)
Counterfactual experience augmented off-policy reinforcement learning
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control
von: Aghabozorgi, Mehran, et al.
Veröffentlicht: (2026)
von: Aghabozorgi, Mehran, et al.
Veröffentlicht: (2026)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
von: Li, Jiaxi, et al.
Veröffentlicht: (2026)
von: Li, Jiaxi, et al.
Veröffentlicht: (2026)
Efficiently Aligning Draft Models via Parameter- and Data-Efficient Adaptation
von: Lin, Luxi, et al.
Veröffentlicht: (2026)
von: Lin, Luxi, et al.
Veröffentlicht: (2026)
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
von: Li, Yuhan, et al.
Veröffentlicht: (2026)
von: Li, Yuhan, et al.
Veröffentlicht: (2026)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
von: Yin, Qingyu, et al.
Veröffentlicht: (2025)
von: Yin, Qingyu, et al.
Veröffentlicht: (2025)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
ParamReL: Learning Parameter Space Representation via Progressively Encoding Bayesian Flow Networks
von: Wu, Zhangkai, et al.
Veröffentlicht: (2024)
von: Wu, Zhangkai, et al.
Veröffentlicht: (2024)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing
von: Pang, Yujuan, et al.
Veröffentlicht: (2026)
von: Pang, Yujuan, et al.
Veröffentlicht: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
Efficient RLVR Training via Weighted Mutual Information Data Selection
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Symbolic Representation for Any-to-Any Generative Tasks
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
von: Huang, Yangyi, et al.
Veröffentlicht: (2026)
von: Huang, Yangyi, et al.
Veröffentlicht: (2026)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
von: Wang, Tao, et al.
Veröffentlicht: (2026)
von: Wang, Tao, et al.
Veröffentlicht: (2026)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
Linear Dynamics in the RLVR Training of Large Language Models
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
von: Lochab, Anamika, et al.
Veröffentlicht: (2026)
von: Lochab, Anamika, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
von: Yin, Qingyu, et al.
Veröffentlicht: (2024) -
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
von: Wang, Hanlin, et al.
Veröffentlicht: (2025) -
Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
von: Xia, Heming, et al.
Veröffentlicht: (2025) -
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
von: Wang, Hanlin, et al.
Veröffentlicht: (2025) -
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
von: Wang, Jian, et al.
Veröffentlicht: (2025)