Pretrain Value, Not Reward: Decoupled Value Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Chenghua, Wang, Lu, Yang, Fangkai, Zhao, Pu, Li, Zhixu, Lin, Qingwei, Zhang, Dongmei, Rajmohan, Saravan, Zhang, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Evolved Reward Learning for LLMs
von: Huang, Chenghua, et al.
Veröffentlicht: (2024)
von: Huang, Chenghua, et al.
Veröffentlicht: (2024)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
von: Wang, Lu, et al.
Veröffentlicht: (2024)
von: Wang, Lu, et al.
Veröffentlicht: (2024)
VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model
von: Zheng, Jiani, et al.
Veröffentlicht: (2025)
von: Zheng, Jiani, et al.
Veröffentlicht: (2025)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
Token-level Proximal Policy Optimization for Query Generation
von: Ouyang, Yichen, et al.
Veröffentlicht: (2024)
von: Ouyang, Yichen, et al.
Veröffentlicht: (2024)
Everything of Thoughts: Defying the Law of Penrose Triangle for Thought Generation
von: Ding, Ruomeng, et al.
Veröffentlicht: (2023)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2023)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
von: Ma, Ming, et al.
Veröffentlicht: (2025)
von: Ma, Ming, et al.
Veröffentlicht: (2025)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
von: Zhang, Jue, et al.
Veröffentlicht: (2025)
von: Zhang, Jue, et al.
Veröffentlicht: (2025)
AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
von: Zhu, Runchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Runchuan, et al.
Veröffentlicht: (2025)
AXIS: Efficient Human-Agent-Computer Interaction with API-First LLM-Based Agents
von: Lu, Junting, et al.
Veröffentlicht: (2024)
von: Lu, Junting, et al.
Veröffentlicht: (2024)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
von: Dong, Hang, et al.
Veröffentlicht: (2024)
von: Dong, Hang, et al.
Veröffentlicht: (2024)
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
Value-Free Policy Optimization via Reward Partitioning
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
von: Fu, Jia, et al.
Veröffentlicht: (2024)
von: Fu, Jia, et al.
Veröffentlicht: (2024)
EfficientRAG: Efficient Retriever for Multi-Hop Question Answering
von: Zhuang, Ziyuan, et al.
Veröffentlicht: (2024)
von: Zhuang, Ziyuan, et al.
Veröffentlicht: (2024)
Beyond State Consistency: Behavior Consistency in Text-Based World Models
von: Huang, Youling, et al.
Veröffentlicht: (2026)
von: Huang, Youling, et al.
Veröffentlicht: (2026)
RuAG: Learned-rule-augmented Generation for Large Language Models
von: Zhang, Yudi, et al.
Veröffentlicht: (2024)
von: Zhang, Yudi, et al.
Veröffentlicht: (2024)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
Reward Models Inherit Value Biases from Pretraining
von: Christian, Brian, et al.
Veröffentlicht: (2026)
von: Christian, Brian, et al.
Veröffentlicht: (2026)
Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented Generation
von: An, Kaikai, et al.
Veröffentlicht: (2024)
von: An, Kaikai, et al.
Veröffentlicht: (2024)
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2024)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
von: Tan, Rongyuan, et al.
Veröffentlicht: (2026)
von: Tan, Rongyuan, et al.
Veröffentlicht: (2026)
Nissist: An Incident Mitigation Copilot based on Troubleshooting Guides
von: An, Kaikai, et al.
Veröffentlicht: (2024)
von: An, Kaikai, et al.
Veröffentlicht: (2024)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
von: Ma, Yunchang, et al.
Veröffentlicht: (2025)
von: Ma, Yunchang, et al.
Veröffentlicht: (2025)
AI Delegates with a Dual Focus: Ensuring Privacy and Strategic Self-Disclosure
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2024)
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
von: Gu, Shangding, et al.
Veröffentlicht: (2024)
Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation
von: Zhang, Xing, et al.
Veröffentlicht: (2025)
von: Zhang, Xing, et al.
Veröffentlicht: (2025)
Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
von: Cheng, Sitao, et al.
Veröffentlicht: (2024)
von: Cheng, Sitao, et al.
Veröffentlicht: (2024)
Quasimetric Value Functions with Dense Rewards
von: Valieva, Khadichabonu, et al.
Veröffentlicht: (2024)
von: Valieva, Khadichabonu, et al.
Veröffentlicht: (2024)
REFA: Reference Free Alignment for multi-preference optimization
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Revisiting Transformer Layer Parameterization Through Causal Energy Minimization
von: Xu, Jin, et al.
Veröffentlicht: (2026)
von: Xu, Jin, et al.
Veröffentlicht: (2026)
Enabling Autonomic Microservice Management through Self-Learning Agents
von: Yu, Fenglin, et al.
Veröffentlicht: (2025)
von: Yu, Fenglin, et al.
Veröffentlicht: (2025)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
Contrastive Learning with Negative Sampling Correction
von: Wang, Lu, et al.
Veröffentlicht: (2024)
von: Wang, Lu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-Evolved Reward Learning for LLMs
von: Huang, Chenghua, et al.
Veröffentlicht: (2024) -
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
von: Wang, Qibin, et al.
Veröffentlicht: (2025) -
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
von: Wang, Lu, et al.
Veröffentlicht: (2024) -
VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model
von: Zheng, Jiani, et al.
Veröffentlicht: (2025) -
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
von: Sun, Lihao, et al.
Veröffentlicht: (2026)