An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
Fuente:
arXiv
Salvato in:
| Autori principali: | Plesner, Andreas, Guzmán, Francisco, Athalye, Anish |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
di: Cai, Xin-Qiang, et al.
Pubblicazione: (2025)
di: Cai, Xin-Qiang, et al.
Pubblicazione: (2025)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
di: He, Haoran, et al.
Pubblicazione: (2025)
di: He, Haoran, et al.
Pubblicazione: (2025)
Rate or Fate? RLV$^\varepsilon$R: Reinforcement Learning with Verifiable Noisy Rewards
di: Rad, Ali, et al.
Pubblicazione: (2026)
di: Rad, Ali, et al.
Pubblicazione: (2026)
Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
di: Islam, Md Mirajul, et al.
Pubblicazione: (2026)
di: Islam, Md Mirajul, et al.
Pubblicazione: (2026)
Learning Robust Reward Machines from Noisy Labels
di: Parac, Roko, et al.
Pubblicazione: (2024)
di: Parac, Roko, et al.
Pubblicazione: (2024)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
di: Lu, Xiaodong, et al.
Pubblicazione: (2026)
di: Lu, Xiaodong, et al.
Pubblicazione: (2026)
Reward Hacking Mitigation using Verifiable Composite Rewards
di: Tarek, Mirza Farhan Bin, et al.
Pubblicazione: (2025)
di: Tarek, Mirza Farhan Bin, et al.
Pubblicazione: (2025)
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
di: Stojanovski, Zafir, et al.
Pubblicazione: (2025)
di: Stojanovski, Zafir, et al.
Pubblicazione: (2025)
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
di: Yoon, Deokgyu, et al.
Pubblicazione: (2026)
di: Yoon, Deokgyu, et al.
Pubblicazione: (2026)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
di: Hu, Haoyu, et al.
Pubblicazione: (2026)
di: Hu, Haoyu, et al.
Pubblicazione: (2026)
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
di: Zhang, Feng, et al.
Pubblicazione: (2026)
di: Zhang, Feng, et al.
Pubblicazione: (2026)
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
di: Song, Kefan, et al.
Pubblicazione: (2025)
di: Song, Kefan, et al.
Pubblicazione: (2025)
Towards Fine-Grained and Verifiable Concept Bottleneck Models
di: Fang, Yingying, et al.
Pubblicazione: (2026)
di: Fang, Yingying, et al.
Pubblicazione: (2026)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
di: Shang, Shuning, et al.
Pubblicazione: (2026)
di: Shang, Shuning, et al.
Pubblicazione: (2026)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
di: Huang, Yu, et al.
Pubblicazione: (2026)
di: Huang, Yu, et al.
Pubblicazione: (2026)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
di: Gunjal, Anisha, et al.
Pubblicazione: (2025)
di: Gunjal, Anisha, et al.
Pubblicazione: (2025)
FLIP Reasoning Challenge
di: Plesner, Andreas, et al.
Pubblicazione: (2025)
di: Plesner, Andreas, et al.
Pubblicazione: (2025)
GOPO: Policy Optimization using Ranked Rewards
di: Choi, Kyuseong, et al.
Pubblicazione: (2026)
di: Choi, Kyuseong, et al.
Pubblicazione: (2026)
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
di: Huang, Jiawei, et al.
Pubblicazione: (2025)
di: Huang, Jiawei, et al.
Pubblicazione: (2025)
Selector-Guided Autonomous Curriculum for One-Shot Reinforcement Learning from Verifiable Rewards
di: Dave, Rudray, et al.
Pubblicazione: (2026)
di: Dave, Rudray, et al.
Pubblicazione: (2026)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
di: Zhang, Zijing, et al.
Pubblicazione: (2025)
di: Zhang, Zijing, et al.
Pubblicazione: (2025)
Optimal Transport for LLM Reward Modeling from Noisy Preference
di: Pan, Licheng, et al.
Pubblicazione: (2026)
di: Pan, Licheng, et al.
Pubblicazione: (2026)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
di: Mansouri, Omar El, et al.
Pubblicazione: (2025)
di: Mansouri, Omar El, et al.
Pubblicazione: (2025)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
di: Helff, Lukas, et al.
Pubblicazione: (2026)
di: Helff, Lukas, et al.
Pubblicazione: (2026)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
di: Roth, Amit, et al.
Pubblicazione: (2026)
di: Roth, Amit, et al.
Pubblicazione: (2026)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2026)
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2026)
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
di: Kim, Yoonjeon, et al.
Pubblicazione: (2025)
di: Kim, Yoonjeon, et al.
Pubblicazione: (2025)
Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation
di: Wang, Longwen, et al.
Pubblicazione: (2026)
di: Wang, Longwen, et al.
Pubblicazione: (2026)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
di: Cho, Dongkyu Derek, et al.
Pubblicazione: (2025)
di: Cho, Dongkyu Derek, et al.
Pubblicazione: (2025)
The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward
di: Li, Long, et al.
Pubblicazione: (2025)
di: Li, Long, et al.
Pubblicazione: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
di: Ackermann, Johannes, et al.
Pubblicazione: (2026)
di: Ackermann, Johannes, et al.
Pubblicazione: (2026)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
di: Wen, Xuexiang, et al.
Pubblicazione: (2026)
di: Wen, Xuexiang, et al.
Pubblicazione: (2026)
DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay
di: Li, Long, et al.
Pubblicazione: (2026)
di: Li, Long, et al.
Pubblicazione: (2026)
Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges
di: Feng, Chen, et al.
Pubblicazione: (2026)
di: Feng, Chen, et al.
Pubblicazione: (2026)
Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards
di: Wang, Zhen, et al.
Pubblicazione: (2025)
di: Wang, Zhen, et al.
Pubblicazione: (2025)
Learning to Explore with Parameter-Space Noise: A Deep Dive into Parameter-Space Noise for Reinforcement Learning with Verifiable Rewards
di: Bai, Bizhe, et al.
Pubblicazione: (2026)
di: Bai, Bizhe, et al.
Pubblicazione: (2026)
Learning to Defer for Causal Discovery with Imperfect Experts
di: Clivio, Oscar, et al.
Pubblicazione: (2025)
di: Clivio, Oscar, et al.
Pubblicazione: (2025)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
di: Ma, Zhengzhao, et al.
Pubblicazione: (2026)
di: Ma, Zhengzhao, et al.
Pubblicazione: (2026)
Is Distance Matrix Enough for Geometric Deep Learning?
di: Li, Zian, et al.
Pubblicazione: (2023)
di: Li, Zian, et al.
Pubblicazione: (2023)
From MNIST to ImageNet: Understanding the Scalability Boundaries of Differentiable Logic Gate Networks
di: Brändle, Sven, et al.
Pubblicazione: (2025)
di: Brändle, Sven, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
di: Cai, Xin-Qiang, et al.
Pubblicazione: (2025) -
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
di: He, Haoran, et al.
Pubblicazione: (2025) -
Rate or Fate? RLV$^\varepsilon$R: Reinforcement Learning with Verifiable Noisy Rewards
di: Rad, Ali, et al.
Pubblicazione: (2026) -
Hierarchical Apprenticeship Learning from Imperfect Demonstrations with Evolving Rewards
di: Islam, Md Mirajul, et al.
Pubblicazione: (2026) -
Learning Robust Reward Machines from Noisy Labels
di: Parac, Roko, et al.
Pubblicazione: (2024)