CoRPO: Adding a Correctness Bias to GRPO Improves Generalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garg, Anisha, Zhang, Claire, Neema, Nishit, Bick, David, Venkatesh, Ganesh, Hestness, Joel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Calibrated Reasoning: An Explanatory Verifier for Dynamic and Efficient Problem-Solving
von: Garg, Anisha, et al.
Veröffentlicht: (2025)
von: Garg, Anisha, et al.
Veröffentlicht: (2025)
The Conductor and the Engine: A Path Towards Co-Designed Reasoning
von: Wang, Yuanxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuanxin, et al.
Veröffentlicht: (2025)
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
von: Neema, Nishit, et al.
Veröffentlicht: (2025)
von: Neema, Nishit, et al.
Veröffentlicht: (2025)
StaRPO: Stability-Augmented Reinforcement Policy Optimization
von: Zhang, Jinghan, et al.
Veröffentlicht: (2026)
von: Zhang, Jinghan, et al.
Veröffentlicht: (2026)
TreeRPO: Tree Relative Policy Optimization
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
von: Zhao, Hanyang, et al.
Veröffentlicht: (2025)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2025)
RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
Scaling with Collapse: Efficient and Predictable Training of LLM Families
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2026)
Don't be lazy: CompleteP enables compute-efficient deep transformers
von: Dey, Nolan, et al.
Veröffentlicht: (2025)
von: Dey, Nolan, et al.
Veröffentlicht: (2025)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
Improving Visual Representation Alignment Generation with GRPO
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
von: Zheng, Jinquan, et al.
Veröffentlicht: (2026)
von: Zheng, Jinquan, et al.
Veröffentlicht: (2026)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Rethinking Refinement: Correcting Generative Bias without Noise Injection
von: Peng, Xin, et al.
Veröffentlicht: (2026)
von: Peng, Xin, et al.
Veröffentlicht: (2026)
What is the Alignment Objective of GRPO?
von: Vojnovic, Milan, et al.
Veröffentlicht: (2025)
von: Vojnovic, Milan, et al.
Veröffentlicht: (2025)
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
von: Bick, Aviv, et al.
Veröffentlicht: (2025)
von: Bick, Aviv, et al.
Veröffentlicht: (2025)
Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2026)
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2026)
EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
von: Block, Adam, et al.
Veröffentlicht: (2025)
von: Block, Adam, et al.
Veröffentlicht: (2025)
Using Early Readouts to Mediate Featural Bias in Distillation
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2023)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2023)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
GRPO is Secretly a Process Reward Model
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
von: Bereket, Michael, et al.
Veröffentlicht: (2025)
von: Bereket, Michael, et al.
Veröffentlicht: (2025)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
Graph Negative Feedback Bias Correction Framework for Adaptive Heterophily Modeling
von: Lv, Jiaqi, et al.
Veröffentlicht: (2026)
von: Lv, Jiaqi, et al.
Veröffentlicht: (2026)
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
von: Yu, Zhiqi, et al.
Veröffentlicht: (2026)
von: Yu, Zhiqi, et al.
Veröffentlicht: (2026)
Adding Conditional Control to Diffusion Models with Reinforcement Learning
von: Zhao, Yulai, et al.
Veröffentlicht: (2024)
von: Zhao, Yulai, et al.
Veröffentlicht: (2024)
GQA-μP: The maximal parameterization update for grouped query attention
von: Chickering, Kyle R., et al.
Veröffentlicht: (2026)
von: Chickering, Kyle R., et al.
Veröffentlicht: (2026)
Transparency and Proportionality in Post-Processing Algorithmic Bias Correction
von: Ferreira, Juliett Suárez, et al.
Veröffentlicht: (2025)
von: Ferreira, Juliett Suárez, et al.
Veröffentlicht: (2025)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
von: Li, Yuming, et al.
Veröffentlicht: (2025)
von: Li, Yuming, et al.
Veröffentlicht: (2025)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
von: Mansouri, Omar El, et al.
Veröffentlicht: (2025)
von: Mansouri, Omar El, et al.
Veröffentlicht: (2025)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
A Unified Framework for Rethinking Policy Divergence Measures in GRPO
von: Wu, Qingyuan, et al.
Veröffentlicht: (2026)
von: Wu, Qingyuan, et al.
Veröffentlicht: (2026)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
von: Wu, Xiongbin, et al.
Veröffentlicht: (2026)
von: Wu, Xiongbin, et al.
Veröffentlicht: (2026)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Calibrated Reasoning: An Explanatory Verifier for Dynamic and Efficient Problem-Solving
von: Garg, Anisha, et al.
Veröffentlicht: (2025) -
The Conductor and the Engine: A Path Towards Co-Designed Reasoning
von: Wang, Yuanxin, et al.
Veröffentlicht: (2025) -
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025) -
From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
von: Neema, Nishit, et al.
Veröffentlicht: (2025) -
StaRPO: Stability-Augmented Reinforcement Policy Optimization
von: Zhang, Jinghan, et al.
Veröffentlicht: (2026)