Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Jiaming, Chen, Longze, Gong, Ze, Chen, Yukun, Wang, Lu, He, Wanwei, Luo, Run, Yang, Min |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
di: Chen, Longze, et al.
Pubblicazione: (2024)
di: Chen, Longze, et al.
Pubblicazione: (2024)
The Unlearnability Phenomenon in RLVR for Language Models
di: Chen, Yulin, et al.
Pubblicazione: (2026)
di: Chen, Yulin, et al.
Pubblicazione: (2026)
Learning Ordinal Probabilistic Reward from Preferences
di: Chen, Longze, et al.
Pubblicazione: (2026)
di: Chen, Longze, et al.
Pubblicazione: (2026)
Self-Distilled RLVR
di: Yang, Chenxu, et al.
Pubblicazione: (2026)
di: Yang, Chenxu, et al.
Pubblicazione: (2026)
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
di: Hong, Joey, et al.
Pubblicazione: (2025)
di: Hong, Joey, et al.
Pubblicazione: (2025)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
di: Zhai, Zepeng, et al.
Pubblicazione: (2026)
di: Zhai, Zepeng, et al.
Pubblicazione: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
di: Huang, Fanding, et al.
Pubblicazione: (2025)
di: Huang, Fanding, et al.
Pubblicazione: (2025)
Linear Dynamics in the RLVR Training of Large Language Models
di: Wang, Tianle, et al.
Pubblicazione: (2026)
di: Wang, Tianle, et al.
Pubblicazione: (2026)
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
di: Wei, Zhepei, et al.
Pubblicazione: (2026)
di: Wei, Zhepei, et al.
Pubblicazione: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
di: He, Bingxiang, et al.
Pubblicazione: (2026)
di: He, Bingxiang, et al.
Pubblicazione: (2026)
Studying the Korean Word-Chain Game with RLVR: Mitigating Reward Conflicts via Curriculum Learning
di: Rho, Donghwan
Pubblicazione: (2025)
di: Rho, Donghwan
Pubblicazione: (2025)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
di: Meng, Haoming, et al.
Pubblicazione: (2026)
di: Meng, Haoming, et al.
Pubblicazione: (2026)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
di: Zhao, Jiale, et al.
Pubblicazione: (2026)
di: Zhao, Jiale, et al.
Pubblicazione: (2026)
Efficient RLVR Training via Weighted Mutual Information Data Selection
di: Zhou, Xinyu, et al.
Pubblicazione: (2026)
di: Zhou, Xinyu, et al.
Pubblicazione: (2026)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
di: Chen, Kun, et al.
Pubblicazione: (2026)
di: Chen, Kun, et al.
Pubblicazione: (2026)
Training Superior Sparse Autoencoders for Instruct Models
di: Li, Jiaming, et al.
Pubblicazione: (2025)
di: Li, Jiaming, et al.
Pubblicazione: (2025)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
di: Cui, Sijia, et al.
Pubblicazione: (2026)
di: Cui, Sijia, et al.
Pubblicazione: (2026)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
Generalization of RLVR Using Causal Reasoning as a Testbed
di: Lu, Brian, et al.
Pubblicazione: (2025)
di: Lu, Brian, et al.
Pubblicazione: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
di: Chen, Peter, et al.
Pubblicazione: (2025)
di: Chen, Peter, et al.
Pubblicazione: (2025)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
di: Kim, Jeonghye, et al.
Pubblicazione: (2026)
di: Kim, Jeonghye, et al.
Pubblicazione: (2026)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation
di: Li, Jiaming, et al.
Pubblicazione: (2025)
di: Li, Jiaming, et al.
Pubblicazione: (2025)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
di: Lochab, Anamika, et al.
Pubblicazione: (2026)
di: Lochab, Anamika, et al.
Pubblicazione: (2026)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
di: Burgess, James, et al.
Pubblicazione: (2026)
di: Burgess, James, et al.
Pubblicazione: (2026)
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
di: Fan, Run-Ze, et al.
Pubblicazione: (2025)
di: Fan, Run-Ze, et al.
Pubblicazione: (2025)
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
di: Xie, Pei-Xi, et al.
Pubblicazione: (2026)
di: Xie, Pei-Xi, et al.
Pubblicazione: (2026)
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
di: Wu, Fang, et al.
Pubblicazione: (2025)
di: Wu, Fang, et al.
Pubblicazione: (2025)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
Anchored Supervised Fine-Tuning
di: Zhu, He, et al.
Pubblicazione: (2025)
di: Zhu, He, et al.
Pubblicazione: (2025)
Bootstrapping Language Models with DPO Implicit Rewards
di: Chen, Changyu, et al.
Pubblicazione: (2024)
di: Chen, Changyu, et al.
Pubblicazione: (2024)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
di: Li, Kaiyuan, et al.
Pubblicazione: (2026)
di: Li, Kaiyuan, et al.
Pubblicazione: (2026)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
Ruler: A Model-Agnostic Method to Control Generated Length for Large Language Models
di: Li, Jiaming, et al.
Pubblicazione: (2024)
di: Li, Jiaming, et al.
Pubblicazione: (2024)
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
di: Mitsuhashi, Ryo, et al.
Pubblicazione: (2026)
di: Mitsuhashi, Ryo, et al.
Pubblicazione: (2026)
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
di: Saif, A F M, et al.
Pubblicazione: (2024)
di: Saif, A F M, et al.
Pubblicazione: (2024)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
di: Choi, Yunho, et al.
Pubblicazione: (2026)
di: Choi, Yunho, et al.
Pubblicazione: (2026)
Huntington Disease Automatic Speech Recognition with Biomarker Supervision
di: Wang, Charles L., et al.
Pubblicazione: (2026)
di: Wang, Charles L., et al.
Pubblicazione: (2026)
Documenti analoghi
-
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
di: Luo, Run, et al.
Pubblicazione: (2025) -
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
di: Luo, Run, et al.
Pubblicazione: (2025) -
Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
di: Chen, Longze, et al.
Pubblicazione: (2024) -
The Unlearnability Phenomenon in RLVR for Language Models
di: Chen, Yulin, et al.
Pubblicazione: (2026) -
Learning Ordinal Probabilistic Reward from Preferences
di: Chen, Longze, et al.
Pubblicazione: (2026)