Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tiwari, Rishabh, Tomar, Aditya, Bamba, Udbhav, Maheswaran, Monishwaran, Yang, Heng, Mahoney, Michael W., Keutzer, Kurt, Gholami, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
by: Maheswaran, Monishwaran, et al.
Published: (2025)
by: Maheswaran, Monishwaran, et al.
Published: (2025)
Squeezed Attention: Accelerating Long Context Length LLM Inference
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
Residual Context Diffusion Language Models
by: Hu, Yuezhou, et al.
Published: (2026)
by: Hu, Yuezhou, et al.
Published: (2026)
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)
by: Tiwari, Rishabh, et al.
Published: (2025)
ETS: Efficient Tree Search for Inference-Time Scaling
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
SciML Agents: Write the Solver, Not the Solution
by: Gaonkar, Saarth, et al.
Published: (2025)
by: Gaonkar, Saarth, et al.
Published: (2025)
AI and Memory Wall
by: Gholami, Amir, et al.
Published: (2024)
by: Gholami, Amir, et al.
Published: (2024)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
by: Subramanian, Shashank, et al.
Published: (2023)
by: Subramanian, Shashank, et al.
Published: (2023)
An LLM Compiler for Parallel Function Calling
by: Kim, Sehoon, et al.
Published: (2023)
by: Kim, Sehoon, et al.
Published: (2023)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
Agentic Test-Time Scaling for WebAgents
by: Lee, Nicholas, et al.
Published: (2026)
by: Lee, Nicholas, et al.
Published: (2026)
SqueezeLLM: Dense-and-Sparse Quantization
by: Kim, Sehoon, et al.
Published: (2023)
by: Kim, Sehoon, et al.
Published: (2023)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
by: Bercovich, Ivan, et al.
Published: (2026)
by: Bercovich, Ivan, et al.
Published: (2026)
TASER: Translation Assessment via Systematic Evaluation and Reasoning
by: Maheswaran, Monishwaran, et al.
Published: (2025)
by: Maheswaran, Monishwaran, et al.
Published: (2025)
Multipole Attention for Efficient Long Context Reasoning
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
by: Lee, Nicholas, et al.
Published: (2024)
by: Lee, Nicholas, et al.
Published: (2024)
The Hackable City
Published: (2020)
Published: (2020)
CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
by: Anand, Neeraj, et al.
Published: (2026)
by: Anand, Neeraj, et al.
Published: (2026)
Characterizing Prompt Compression Methods for Long Context Inference
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
RRM: Robust Reward Model Training Mitigates Reward Hacking
by: Liu, Tianqi, et al.
Published: (2024)
by: Liu, Tianqi, et al.
Published: (2024)
Reward Centering
by: Naik, Abhishek, et al.
Published: (2024)
by: Naik, Abhishek, et al.
Published: (2024)
Are PPO-ed Language Models Hackable?
by: Anand, Suraj, et al.
Published: (2024)
by: Anand, Suraj, et al.
Published: (2024)
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
by: Maheswaran, Monishwaran, et al.
Published: (2026)
by: Maheswaran, Monishwaran, et al.
Published: (2026)
$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners
by: Singh, Harman, et al.
Published: (2026)
by: Singh, Harman, et al.
Published: (2026)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
by: Hooper, Coleman, et al.
Published: (2026)
by: Hooper, Coleman, et al.
Published: (2026)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
by: Pala, Tej Deep, et al.
Published: (2025)
by: Pala, Tej Deep, et al.
Published: (2025)
Reward-Preserving Attacks For Robust Reinforcement Learning
by: Schott, Lucas, et al.
Published: (2026)
by: Schott, Lucas, et al.
Published: (2026)
Concentration of Cumulative Reward in Markov Decision Processes
by: Sayedana, Borna, et al.
Published: (2024)
by: Sayedana, Borna, et al.
Published: (2024)
CDLM: Consistency Diffusion Language Models For Faster Sampling
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
SPEED: Speculative Pipelined Execution for Efficient Decoding
by: Hooper, Coleman, et al.
Published: (2023)
by: Hooper, Coleman, et al.
Published: (2023)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
by: Gureja, Srishti, et al.
Published: (2024)
by: Gureja, Srishti, et al.
Published: (2024)
Adversarial Training for Process Reward Models
by: Juneja, Gurusha, et al.
Published: (2025)
by: Juneja, Gurusha, et al.
Published: (2025)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
DOT-MoE: Differentiable Optimal Transport for MoEfication
by: Bamba, Udbhav, et al.
Published: (2026)
by: Bamba, Udbhav, et al.
Published: (2026)
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
by: Bamba, Udbhav, et al.
Published: (2025)
by: Bamba, Udbhav, et al.
Published: (2025)
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
by: Qiu, Zhisong, et al.
Published: (2026)
by: Qiu, Zhisong, et al.
Published: (2026)
Similar Items
-
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
by: Maheswaran, Monishwaran, et al.
Published: (2025) -
Squeezed Attention: Accelerating Long Context Length LLM Inference
by: Hooper, Coleman, et al.
Published: (2024) -
Residual Context Diffusion Language Models
by: Hu, Yuezhou, et al.
Published: (2026) -
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
by: Tomar, Aditya, et al.
Published: (2025) -
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)