Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
Fuente:
arXiv
Guardado en:
| Autores principales: | Khalifa, Muhammad, Khan, Zohaib, Tafveez, Omer, Peng, Hao, Wang, Lu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generalization of RLVR Using Causal Reasoning as a Testbed
por: Lu, Brian, et al.
Publicado: (2025)
por: Lu, Brian, et al.
Publicado: (2025)
Plasticity vs. Rigidity: The Impact of Low-Rank Adapters on Reasoning on a Micro-Budget
por: Khan, Zohaib, et al.
Publicado: (2026)
por: Khan, Zohaib, et al.
Publicado: (2026)
Process Reward Models That Think
por: Khalifa, Muhammad, et al.
Publicado: (2025)
por: Khalifa, Muhammad, et al.
Publicado: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
por: Chen, Peter, et al.
Publicado: (2025)
por: Chen, Peter, et al.
Publicado: (2025)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
por: Helff, Lukas, et al.
Publicado: (2026)
por: Helff, Lukas, et al.
Publicado: (2026)
RLPR: Extrapolating RLVR to General Domains without Verifiers
por: Yu, Tianyu, et al.
Publicado: (2025)
por: Yu, Tianyu, et al.
Publicado: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
por: Duo, Jiangshan, et al.
Publicado: (2026)
por: Duo, Jiangshan, et al.
Publicado: (2026)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
por: Cui, Sijia, et al.
Publicado: (2026)
por: Cui, Sijia, et al.
Publicado: (2026)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
por: Sahoo, Subramanyam
Publicado: (2026)
por: Sahoo, Subramanyam
Publicado: (2026)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
por: Chen, Kun, et al.
Publicado: (2026)
por: Chen, Kun, et al.
Publicado: (2026)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
por: Gu, Hengrui, et al.
Publicado: (2026)
por: Gu, Hengrui, et al.
Publicado: (2026)
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
por: Wu, Fang, et al.
Publicado: (2025)
por: Wu, Fang, et al.
Publicado: (2025)
TracrBench: Generating Interpretability Testbeds with Large Language Models
por: Thurnherr, Hannes, et al.
Publicado: (2024)
por: Thurnherr, Hannes, et al.
Publicado: (2024)
Multi-Turn Code Generation Through Single-Step Rewards
por: Jain, Arnav Kumar, et al.
Publicado: (2025)
por: Jain, Arnav Kumar, et al.
Publicado: (2025)
On Teacher Hacking in Language Model Distillation
por: Tiapkin, Daniil, et al.
Publicado: (2025)
por: Tiapkin, Daniil, et al.
Publicado: (2025)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
por: Lee, Chanuk, et al.
Publicado: (2026)
por: Lee, Chanuk, et al.
Publicado: (2026)
Scaling Unverifiable Rewards: A Case Study on Visual Insights
por: Gan, Shuyu, et al.
Publicado: (2025)
por: Gan, Shuyu, et al.
Publicado: (2025)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
por: Chen, Zhipeng, et al.
Publicado: (2026)
por: Chen, Zhipeng, et al.
Publicado: (2026)
Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention
por: Khan, Zohaib, et al.
Publicado: (2024)
por: Khan, Zohaib, et al.
Publicado: (2024)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
por: Fan, Lishui, et al.
Publicado: (2025)
por: Fan, Lishui, et al.
Publicado: (2025)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
por: Meng, Haoming, et al.
Publicado: (2026)
por: Meng, Haoming, et al.
Publicado: (2026)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
por: Yu, Zhuohao, et al.
Publicado: (2024)
por: Yu, Zhuohao, et al.
Publicado: (2024)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
por: Peng, Miao, et al.
Publicado: (2025)
por: Peng, Miao, et al.
Publicado: (2025)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
por: Casademunt, Helena, et al.
Publicado: (2026)
por: Casademunt, Helena, et al.
Publicado: (2026)
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
por: Zhang, Ruiyi, et al.
Publicado: (2025)
por: Zhang, Ruiyi, et al.
Publicado: (2025)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
por: Kim, Soeun, et al.
Publicado: (2026)
por: Kim, Soeun, et al.
Publicado: (2026)
AgentRM: Enhancing Agent Generalization with Reward Modeling
por: Xia, Yu, et al.
Publicado: (2025)
por: Xia, Yu, et al.
Publicado: (2025)
A State-of-the-Art SQL Reasoning Model using RLVR
por: Ali, Alnur, et al.
Publicado: (2025)
por: Ali, Alnur, et al.
Publicado: (2025)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
por: Ye, Hao, et al.
Publicado: (2026)
por: Ye, Hao, et al.
Publicado: (2026)
Cloning Ideology and Style using Deep Learning
por: Beg, Omer, et al.
Publicado: (2022)
por: Beg, Omer, et al.
Publicado: (2022)
$\mathbf{(N,K)}$-Puzzle: A Cost-Efficient Testbed for Benchmarking Reinforcement Learning Algorithms in Generative Language Model
por: Zhang, Yufeng, et al.
Publicado: (2024)
por: Zhang, Yufeng, et al.
Publicado: (2024)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
por: Burgess, James, et al.
Publicado: (2026)
por: Burgess, James, et al.
Publicado: (2026)
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
por: Xie, Pei-Xi, et al.
Publicado: (2026)
por: Xie, Pei-Xi, et al.
Publicado: (2026)
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
por: Song, Kefan, et al.
Publicado: (2025)
por: Song, Kefan, et al.
Publicado: (2025)
Ejemplares similares
-
Generalization of RLVR Using Causal Reasoning as a Testbed
por: Lu, Brian, et al.
Publicado: (2025) -
Plasticity vs. Rigidity: The Impact of Low-Rank Adapters on Reasoning on a Micro-Budget
por: Khan, Zohaib, et al.
Publicado: (2026) -
Process Reward Models That Think
por: Khalifa, Muhammad, et al.
Publicado: (2025) -
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025) -
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)