Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Longwen, Liu, Yirui, Wu, Xuan'er, Hu, Xiaohui, Fan, Yuankai, Yu, Kaidong, Weng, Qizhen, Xi, Wei, Li, Xuelong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
por: Fan, Yuankai, et al.
Publicado: (2025)
por: Fan, Yuankai, et al.
Publicado: (2025)
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
por: Liang, Yuanzhi, et al.
Publicado: (2026)
por: Liang, Yuanzhi, et al.
Publicado: (2026)
Prompt-Level Reward Specifications for Open-Ended Post-Training
por: Weng, Zijun, et al.
Publicado: (2026)
por: Weng, Zijun, et al.
Publicado: (2026)
PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
por: Zhang, Yang, et al.
Publicado: (2026)
por: Zhang, Yang, et al.
Publicado: (2026)
Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression
por: Qi, Ruoling, et al.
Publicado: (2026)
por: Qi, Ruoling, et al.
Publicado: (2026)
Improve LLM-as-a-Judge Ability as a General Ability
por: Yu, Jiachen, et al.
Publicado: (2025)
por: Yu, Jiachen, et al.
Publicado: (2025)
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
por: He, Yinghui, et al.
Publicado: (2026)
por: He, Yinghui, et al.
Publicado: (2026)
Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards
por: Jolfaei, Erfan Aghadavoodi, et al.
Publicado: (2026)
por: Jolfaei, Erfan Aghadavoodi, et al.
Publicado: (2026)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025)
por: Gunjal, Anisha, et al.
Publicado: (2025)
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
por: Mroueh, Youssef
Publicado: (2025)
por: Mroueh, Youssef
Publicado: (2025)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
por: Fan, Lishui, et al.
Publicado: (2025)
por: Fan, Lishui, et al.
Publicado: (2025)
SecureCodeRL: Security-Aware Reinforcement Learning for Code Generation with Partial-Credit Rewards
por: Sijwali, Suryansh Singh, et al.
Publicado: (2026)
por: Sijwali, Suryansh Singh, et al.
Publicado: (2026)
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering
por: Fan, Shicheng, et al.
Publicado: (2026)
por: Fan, Shicheng, et al.
Publicado: (2026)
DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation
por: Zhu, Speed, et al.
Publicado: (2025)
por: Zhu, Speed, et al.
Publicado: (2025)
Video Models Can Reason with Verifiable Rewards
por: Zhu, Tinghui, et al.
Publicado: (2026)
por: Zhu, Tinghui, et al.
Publicado: (2026)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
por: Tang, Yunhao, et al.
Publicado: (2025)
por: Tang, Yunhao, et al.
Publicado: (2025)
Multi-Turn Code Generation Through Single-Step Rewards
por: Jain, Arnav Kumar, et al.
Publicado: (2025)
por: Jain, Arnav Kumar, et al.
Publicado: (2025)
On Reward Transferability in Adversarial Inverse Reinforcement Learning: Insights from Random Matrix Theory
por: Zhang, Yangchun, et al.
Publicado: (2024)
por: Zhang, Yangchun, et al.
Publicado: (2024)
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
por: Jiang, Yuxin, et al.
Publicado: (2026)
por: Jiang, Yuxin, et al.
Publicado: (2026)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
por: Huang, Jiawei, et al.
Publicado: (2026)
por: Huang, Jiawei, et al.
Publicado: (2026)
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
por: Zhang, Feng, et al.
Publicado: (2026)
por: Zhang, Feng, et al.
Publicado: (2026)
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
por: Jiang, Yuhua, et al.
Publicado: (2025)
por: Jiang, Yuhua, et al.
Publicado: (2025)
Estimating the Joint Distribution of Two Binary Variables with Marginal Statistics
por: Shang, Longwen, et al.
Publicado: (2025)
por: Shang, Longwen, et al.
Publicado: (2025)
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
por: Cai, Xin-Qiang, et al.
Publicado: (2025)
por: Cai, Xin-Qiang, et al.
Publicado: (2025)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
por: Hu, Haoyu, et al.
Publicado: (2026)
por: Hu, Haoyu, et al.
Publicado: (2026)
Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA Generation
por: Zou, Qingyun, et al.
Publicado: (2026)
por: Zou, Qingyun, et al.
Publicado: (2026)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
por: Hu, Wentao, et al.
Publicado: (2026)
por: Hu, Wentao, et al.
Publicado: (2026)
DLA-Count: Dynamic Label Assignment Network for Dense Cell Distribution Counting
por: Yan, Yuqing, et al.
Publicado: (2025)
por: Yan, Yuqing, et al.
Publicado: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
por: Liu, Xiaoyuan, et al.
Publicado: (2025)
por: Liu, Xiaoyuan, et al.
Publicado: (2025)
Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards
por: Lara, Luis, et al.
Publicado: (2026)
por: Lara, Luis, et al.
Publicado: (2026)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
por: Zhang, Xin, et al.
Publicado: (2026)
por: Zhang, Xin, et al.
Publicado: (2026)
Beyond Imitation: Recovering Dense Rewards from Demonstrations
por: Li, Jiangnan, et al.
Publicado: (2025)
por: Li, Jiangnan, et al.
Publicado: (2025)
Driving Beyond Privilege: Distilling Dense-Reward Knowledge into Sparse-Reward Policies
por: Khanzada, Feeza Khan, et al.
Publicado: (2025)
por: Khanzada, Feeza Khan, et al.
Publicado: (2025)
Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
por: Bo, Zitong, et al.
Publicado: (2025)
por: Bo, Zitong, et al.
Publicado: (2025)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
por: Liu, Yule, et al.
Publicado: (2025)
por: Liu, Yule, et al.
Publicado: (2025)
Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards
por: Tang, Xinyu, et al.
Publicado: (2025)
por: Tang, Xinyu, et al.
Publicado: (2025)
Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards
por: Heckel, Reinhard, et al.
Publicado: (2026)
por: Heckel, Reinhard, et al.
Publicado: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
por: Lu, Xiaodong, et al.
Publicado: (2026)
por: Lu, Xiaodong, et al.
Publicado: (2026)
Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards
por: Zhu, Yuxuan, et al.
Publicado: (2026)
por: Zhu, Yuxuan, et al.
Publicado: (2026)
Rethinking Adversarial Inverse Reinforcement Learning: Policy Imitation, Transferable Reward Recovery and Algebraic Equilibrium Proof
por: Zhang, Yangchun, et al.
Publicado: (2024)
por: Zhang, Yangchun, et al.
Publicado: (2024)
Ejemplares similares
-
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
por: Fan, Yuankai, et al.
Publicado: (2025) -
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
por: Liang, Yuanzhi, et al.
Publicado: (2026) -
Prompt-Level Reward Specifications for Open-Ended Post-Training
por: Weng, Zijun, et al.
Publicado: (2026) -
PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
por: Zhang, Yang, et al.
Publicado: (2026) -
Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression
por: Qi, Ruoling, et al.
Publicado: (2026)