Improving LLM-Generated Code Quality with GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Robeyns, Maxime, Aitchison, Laurence |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Self-Improving Coding Agent
by: Robeyns, Maxime, et al.
Published: (2025)
by: Robeyns, Maxime, et al.
Published: (2025)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025)
by: Bowyer, Sam, et al.
Published: (2025)
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
iGRPO: Self-Feedback-Driven LLM Reasoning
by: Hatamizadeh, Ali, et al.
Published: (2026)
by: Hatamizadeh, Ali, et al.
Published: (2026)
CoRPO: Adding a Correctness Bias to GRPO Improves Generalization
by: Garg, Anisha, et al.
Published: (2025)
by: Garg, Anisha, et al.
Published: (2025)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
by: Akli, Amal, et al.
Published: (2026)
by: Akli, Amal, et al.
Published: (2026)
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
by: Ekbote, Chanakya, et al.
Published: (2025)
by: Ekbote, Chanakya, et al.
Published: (2025)
Bayesian Reward Models for LLM Alignment
by: Yang, Adam X., et al.
Published: (2024)
by: Yang, Adam X., et al.
Published: (2024)
Rodent-Bench
by: Heap, Thomas, et al.
Published: (2026)
by: Heap, Thomas, et al.
Published: (2026)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
by: Parthasarathi, Prasanna, et al.
Published: (2025)
by: Parthasarathi, Prasanna, et al.
Published: (2025)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
by: Chen, Minghan, et al.
Published: (2025)
by: Chen, Minghan, et al.
Published: (2025)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
by: Pennino, Federico, et al.
Published: (2025)
by: Pennino, Federico, et al.
Published: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training
by: Ai, Rui, et al.
Published: (2026)
by: Ai, Rui, et al.
Published: (2026)
Memorize or Generalize? Evaluating LLM Code Generation with Code Rewriting
by: Zhang, Lizhe, et al.
Published: (2025)
by: Zhang, Lizhe, et al.
Published: (2025)
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
by: Zhang, Xiaoying, et al.
Published: (2025)
by: Zhang, Xiaoying, et al.
Published: (2025)
ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
by: Zhao, Yicong, et al.
Published: (2025)
by: Zhao, Yicong, et al.
Published: (2025)
Optimizing for Persuasion Improves LLM Generalization: Evidence from Quality-Diversity Evolution of Debate Strategies
by: Reedi, Aksel Joonas, et al.
Published: (2025)
by: Reedi, Aksel Joonas, et al.
Published: (2025)
Instruction Tuning With Loss Over Instructions
by: Shi, Zhengyan, et al.
Published: (2024)
by: Shi, Zhengyan, et al.
Published: (2024)
Comparative Analysis and Parametric Tuning of PPO, GRPO, and DAPO for LLM Reasoning Enhancement
by: Lian, Yongsheng
Published: (2025)
by: Lian, Yongsheng
Published: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning
by: Yin, Shouyu, et al.
Published: (2026)
by: Yin, Shouyu, et al.
Published: (2026)
Planning In Natural Language Improves LLM Search For Code Generation
by: Wang, Evan, et al.
Published: (2024)
by: Wang, Evan, et al.
Published: (2024)
Quality Assurance of LLM-generated Code: Addressing Non-Functional Quality Characteristics
by: Sun, Xin, et al.
Published: (2025)
by: Sun, Xin, et al.
Published: (2025)
What is the Alignment Objective of GRPO?
by: Vojnovic, Milan, et al.
Published: (2025)
by: Vojnovic, Milan, et al.
Published: (2025)
AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
by: Ahtisham, Bakhtawar, et al.
Published: (2025)
by: Ahtisham, Bakhtawar, et al.
Published: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
On Iterative Evaluation and Enhancement of Code Quality Using GPT-4o
by: Liu, Rundong, et al.
Published: (2025)
by: Liu, Rundong, et al.
Published: (2025)
Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code
by: Blain, Dominik, et al.
Published: (2026)
by: Blain, Dominik, et al.
Published: (2026)
CodeGuard: Improving LLM Guardrails in CS Education
by: Raihan, Nishat, et al.
Published: (2026)
by: Raihan, Nishat, et al.
Published: (2026)
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Beyond Retrieval: Improving Evidence Quality for LLM-based Multimodal Fact-Checking
by: Ou, Haoran, et al.
Published: (2025)
by: Ou, Haoran, et al.
Published: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
Similar Items
-
A Self-Improving Coding Agent
by: Robeyns, Maxime, et al.
Published: (2025) -
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024) -
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025) -
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023) -
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair
by: Li, Jia, et al.
Published: (2026)