PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Ruiyi, Lv, Yongxuan, Meng, Xianhui, Song, Jiahe, Wang, Chao, Jiang, Chen, Cheng, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
by: Zhou, Yangyang, et al.
Published: (2026)
by: Zhou, Yangyang, et al.
Published: (2026)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
by: Han, Xudong, et al.
Published: (2025)
by: Han, Xudong, et al.
Published: (2025)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
by: Saxena, Udit
Published: (2025)
by: Saxena, Udit
Published: (2025)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
by: Sun, Yuhui, et al.
Published: (2025)
by: Sun, Yuhui, et al.
Published: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)
by: Kiashemshaki, Kiana, et al.
Published: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
by: Schesch, Benedikt, et al.
Published: (2026)
by: Schesch, Benedikt, et al.
Published: (2026)
Survey Transfer Learning: Recycling Data with Silicon Responses
by: Amini, Ali
Published: (2025)
by: Amini, Ali
Published: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
by: Liu, Zhongxin, et al.
Published: (2025)
by: Liu, Zhongxin, et al.
Published: (2025)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026)
by: Hilasaca, Kenji, et al.
Published: (2026)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
by: Lan, Guangchen, et al.
Published: (2026)
by: Lan, Guangchen, et al.
Published: (2026)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
by: He, Yanjin, et al.
Published: (2025)
by: He, Yanjin, et al.
Published: (2025)
Reflective Verbal Reward Design for Pluralistic Alignment
by: Blair, Carter, et al.
Published: (2025)
by: Blair, Carter, et al.
Published: (2025)
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
by: Gwak, Jiho, et al.
Published: (2025)
by: Gwak, Jiho, et al.
Published: (2025)
On the Influence of Discourse Relations in Persuasive Texts
by: Turk, Nawar, et al.
Published: (2025)
by: Turk, Nawar, et al.
Published: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
by: Easley, Eric, et al.
Published: (2026)
by: Easley, Eric, et al.
Published: (2026)
Calibrated Confidence Estimation for Tabular Question Answering
by: Voss, Lukas
Published: (2026)
by: Voss, Lukas
Published: (2026)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
by: Liao, Jianxing, et al.
Published: (2025)
by: Liao, Jianxing, et al.
Published: (2025)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
by: Chen, Xinjie, et al.
Published: (2026)
by: Chen, Xinjie, et al.
Published: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026)
by: McCann, Jordan F.
Published: (2026)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
by: Srishty, Sharmin Sultana, et al.
Published: (2026)
by: Srishty, Sharmin Sultana, et al.
Published: (2026)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
by: Ye, Hua, et al.
Published: (2025)
by: Ye, Hua, et al.
Published: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
by: Kugler, Kai
Published: (2025)
by: Kugler, Kai
Published: (2025)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
by: Resck, Lucas, et al.
Published: (2026)
by: Resck, Lucas, et al.
Published: (2026)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
by: Rozonoyer, Benjamin, et al.
Published: (2026)
by: Rozonoyer, Benjamin, et al.
Published: (2026)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
by: Bertina, Abbas, et al.
Published: (2025)
by: Bertina, Abbas, et al.
Published: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
by: Bakish, Yarden, et al.
Published: (2025)
by: Bakish, Yarden, et al.
Published: (2025)
QuAnTS: Question Answering on Time Series
by: Divo, Felix, et al.
Published: (2025)
by: Divo, Felix, et al.
Published: (2025)
Similar Items
-
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
by: Zhou, Yangyang, et al.
Published: (2026) -
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
by: Nakamura, Mason, et al.
Published: (2025) -
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
by: Zhang, Bowen, et al.
Published: (2025) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025) -
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
by: Zhang, Yizhou, et al.
Published: (2025)