Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Shan, Zikang, Zhong, Han, Wang, Liwei, Zhao, Li |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DPO Meets PPO: Reinforced Token Optimization for RLHF
di: Zhong, Han, et al.
Pubblicazione: (2024)
di: Zhong, Han, et al.
Pubblicazione: (2024)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
di: Choi, Yunho, et al.
Pubblicazione: (2026)
di: Choi, Yunho, et al.
Pubblicazione: (2026)
Generative Value Conflicts Reveal LLM Priorities
di: Liu, Andy, et al.
Pubblicazione: (2025)
di: Liu, Andy, et al.
Pubblicazione: (2025)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
di: Jin, Haoran, et al.
Pubblicazione: (2025)
di: Jin, Haoran, et al.
Pubblicazione: (2025)
Value Augmented Sampling for Language Model Alignment and Personalization
di: Han, Seungwook, et al.
Pubblicazione: (2024)
di: Han, Seungwook, et al.
Pubblicazione: (2024)
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning
di: Chen, Zizhe, et al.
Pubblicazione: (2026)
di: Chen, Zizhe, et al.
Pubblicazione: (2026)
FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
di: Zhong, Shan, et al.
Pubblicazione: (2025)
di: Zhong, Shan, et al.
Pubblicazione: (2025)
Relative Value Biases in Large Language Models
di: Hayes, William M., et al.
Pubblicazione: (2024)
di: Hayes, William M., et al.
Pubblicazione: (2024)
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
di: Zhong, Han, et al.
Pubblicazione: (2025)
di: Zhong, Han, et al.
Pubblicazione: (2025)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
di: Kang, Yipeng, et al.
Pubblicazione: (2024)
di: Kang, Yipeng, et al.
Pubblicazione: (2024)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
di: Huang, Saffron, et al.
Pubblicazione: (2025)
di: Huang, Saffron, et al.
Pubblicazione: (2025)
SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning
di: Ai, Zhengyang, et al.
Pubblicazione: (2026)
di: Ai, Zhengyang, et al.
Pubblicazione: (2026)
Value-Aware Numerical Representations for Transformer Language Models
di: Dutulescu, Andreea, et al.
Pubblicazione: (2026)
di: Dutulescu, Andreea, et al.
Pubblicazione: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
$V_0$: A Generalist Value Model for Any Policy at State Zero
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2026)
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2026)
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2026)
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2026)
CARL: Criticality-Aware Agentic Reinforcement Learning
di: Shen, Leyang, et al.
Pubblicazione: (2025)
di: Shen, Leyang, et al.
Pubblicazione: (2025)
CovidLLM: A Robust Large Language Model with Missing Value Adaptation and Multi-Objective Learning Strategy for Predicting Disease Severity and Clinical Outcomes in COVID-19 Patients
di: Zhu, Shengjun, et al.
Pubblicazione: (2024)
di: Zhu, Shengjun, et al.
Pubblicazione: (2024)
KaSA: Knowledge-Aware Singular-Value Adaptation of Large Language Models
di: Wang, Fan, et al.
Pubblicazione: (2024)
di: Wang, Fan, et al.
Pubblicazione: (2024)
Explaining Large Language Models Decisions Using Shapley Values
di: Mohammadi, Behnam
Pubblicazione: (2024)
di: Mohammadi, Behnam
Pubblicazione: (2024)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
di: Wang, Xinyi, et al.
Pubblicazione: (2024)
di: Wang, Xinyi, et al.
Pubblicazione: (2024)
VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
di: Zhu, Dingwei, et al.
Pubblicazione: (2025)
di: Zhu, Dingwei, et al.
Pubblicazione: (2025)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
di: Rammal, Mohamad Rida, et al.
Pubblicazione: (2024)
di: Rammal, Mohamad Rida, et al.
Pubblicazione: (2024)
Regurgitative Training: The Value of Real Data in Training Large Language Models
di: Zhang, Jinghui, et al.
Pubblicazione: (2024)
di: Zhang, Jinghui, et al.
Pubblicazione: (2024)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
di: Wu, Mian, et al.
Pubblicazione: (2025)
di: Wu, Mian, et al.
Pubblicazione: (2025)
Learning to Reason from Feedback at Test-Time
di: Li, Yanyang, et al.
Pubblicazione: (2025)
di: Li, Yanyang, et al.
Pubblicazione: (2025)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
di: Zheng, Shenyan, et al.
Pubblicazione: (2026)
di: Zheng, Shenyan, et al.
Pubblicazione: (2026)
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
di: Zhong, Tianle, et al.
Pubblicazione: (2026)
di: Zhong, Tianle, et al.
Pubblicazione: (2026)
Reward Models Inherit Value Biases from Pretraining
di: Christian, Brian, et al.
Pubblicazione: (2026)
di: Christian, Brian, et al.
Pubblicazione: (2026)
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
di: Xu, Wanghan, et al.
Pubblicazione: (2026)
di: Xu, Wanghan, et al.
Pubblicazione: (2026)
Language Models can Self-Improve at State-Value Estimation for Better Search
di: Mendes, Ethan, et al.
Pubblicazione: (2025)
di: Mendes, Ethan, et al.
Pubblicazione: (2025)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
di: Zhang, Wenbo, et al.
Pubblicazione: (2025)
di: Zhang, Wenbo, et al.
Pubblicazione: (2025)
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
di: Wang, Zhaoyang, et al.
Pubblicazione: (2026)
di: Wang, Zhaoyang, et al.
Pubblicazione: (2026)
Learning a Generative Meta-Model of LLM Activations
di: Luo, Grace, et al.
Pubblicazione: (2026)
di: Luo, Grace, et al.
Pubblicazione: (2026)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
di: Yue, Yuxuan, et al.
Pubblicazione: (2024)
di: Yue, Yuxuan, et al.
Pubblicazione: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
di: Liu, Qihao, et al.
Pubblicazione: (2025)
di: Liu, Qihao, et al.
Pubblicazione: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
di: Wang, Boxin, et al.
Pubblicazione: (2025)
di: Wang, Boxin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DPO Meets PPO: Reinforced Token Optimization for RLHF
di: Zhong, Han, et al.
Pubblicazione: (2024) -
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
di: Choi, Yunho, et al.
Pubblicazione: (2026) -
Generative Value Conflicts Reveal LLM Priorities
di: Liu, Andy, et al.
Pubblicazione: (2025) -
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024) -
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
di: Jin, Haoran, et al.
Pubblicazione: (2025)