When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Rui, Tang, Ruixiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
von: Deng, Wenlong, et al.
Veröffentlicht: (2026)
von: Deng, Wenlong, et al.
Veröffentlicht: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
von: Wang, Songtao, et al.
Veröffentlicht: (2026)
von: Wang, Songtao, et al.
Veröffentlicht: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
A Representation-Level Assessment of Bias Mitigation in Foundation Models
von: Nizhnichenkov, Svetoslav, et al.
Veröffentlicht: (2026)
von: Nizhnichenkov, Svetoslav, et al.
Veröffentlicht: (2026)
Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
von: Wu, Rui, et al.
Veröffentlicht: (2025)
von: Wu, Rui, et al.
Veröffentlicht: (2025)
DBR: Divergence-Based Regularization for Debiasing Natural Language Understanding Models
von: Li, Zihao, et al.
Veröffentlicht: (2025)
von: Li, Zihao, et al.
Veröffentlicht: (2025)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
Scalable Ensembling For Mitigating Reward Overoptimisation
von: Ahmed, Ahmed M., et al.
Veröffentlicht: (2024)
von: Ahmed, Ahmed M., et al.
Veröffentlicht: (2024)
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
von: Xie, Tong, et al.
Veröffentlicht: (2025)
von: Xie, Tong, et al.
Veröffentlicht: (2025)
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
von: Tao, Leitian, et al.
Veröffentlicht: (2025)
von: Tao, Leitian, et al.
Veröffentlicht: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
von: Tan, Chuyi, et al.
Veröffentlicht: (2025)
von: Tan, Chuyi, et al.
Veröffentlicht: (2025)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
von: Lu, Keming, et al.
Veröffentlicht: (2024)
von: Lu, Keming, et al.
Veröffentlicht: (2024)
RRM: Robust Reward Model Training Mitigates Reward Hacking
von: Liu, Tianqi, et al.
Veröffentlicht: (2024)
von: Liu, Tianqi, et al.
Veröffentlicht: (2024)
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use
von: Kumar, Abhijit, et al.
Veröffentlicht: (2026)
von: Kumar, Abhijit, et al.
Veröffentlicht: (2026)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
von: Lian, Jiesong, et al.
Veröffentlicht: (2025)
von: Lian, Jiesong, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2026)
von: Yu, Yongcan, et al.
Veröffentlicht: (2026)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
von: Jang, Eyon, et al.
Veröffentlicht: (2026)
von: Jang, Eyon, et al.
Veröffentlicht: (2026)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
von: Chen, Tong, et al.
Veröffentlicht: (2025)
von: Chen, Tong, et al.
Veröffentlicht: (2025)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
von: Ma, Qiyao, et al.
Veröffentlicht: (2026)
von: Ma, Qiyao, et al.
Veröffentlicht: (2026)
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
von: Wang, Li, et al.
Veröffentlicht: (2026)
von: Wang, Li, et al.
Veröffentlicht: (2026)
When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
On Teacher Hacking in Language Model Distillation
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
von: Li, Jiaqian, et al.
Veröffentlicht: (2026)
von: Li, Jiaqian, et al.
Veröffentlicht: (2026)
Demystifying When Pruning Works via Representation Hierarchies
von: He, Shwai, et al.
Veröffentlicht: (2026)
von: He, Shwai, et al.
Veröffentlicht: (2026)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
von: Le, Khoi, et al.
Veröffentlicht: (2026)
von: Le, Khoi, et al.
Veröffentlicht: (2026)
Studying the Korean Word-Chain Game with RLVR: Mitigating Reward Conflicts via Curriculum Learning
von: Rho, Donghwan
Veröffentlicht: (2025)
von: Rho, Donghwan
Veröffentlicht: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
von: Tarek, Mirza Farhan Bin, et al.
Veröffentlicht: (2025)
von: Tarek, Mirza Farhan Bin, et al.
Veröffentlicht: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
TLDR: Token-Level Detective Reward Model for Large Vision Language Models
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
von: Cao, Shuirong, et al.
Veröffentlicht: (2024)
von: Cao, Shuirong, et al.
Veröffentlicht: (2024)
Bayesian Preference Learning for Test-Time Steerable Reward Models
von: Hong, Jiwoo, et al.
Veröffentlicht: (2026)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2026)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025) -
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024) -
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
von: Deng, Wenlong, et al.
Veröffentlicht: (2026) -
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026) -
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)