Detecting and Suppressing Reward Hacking with Gradient Fingerprints
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Songtao, Pham, Quang Hieu, Yin, Fangcong, Wang, Xinpeng, Chen, Jocelyn Qiaochu, Durrett, Greg, Ye, Xi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LoFiT: Localized Fine-tuning on LLM Representations
von: Yin, Fangcong, et al.
Veröffentlicht: (2024)
von: Yin, Fangcong, et al.
Veröffentlicht: (2024)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
von: Xin, Yutong, et al.
Veröffentlicht: (2026)
von: Xin, Yutong, et al.
Veröffentlicht: (2026)
Understanding Synthetic Context Extension via Retrieval Heads
von: Zhao, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2024)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2024)
von: Sprague, Zayne, et al.
Veröffentlicht: (2024)
CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation
von: Khatry, Anirudh, et al.
Veröffentlicht: (2025)
von: Khatry, Anirudh, et al.
Veröffentlicht: (2025)
Learning Composable Chains-of-Thought
von: Yin, Fangcong, et al.
Veröffentlicht: (2025)
von: Yin, Fangcong, et al.
Veröffentlicht: (2025)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
von: Divekar, Abhishek, et al.
Veröffentlicht: (2024)
von: Divekar, Abhishek, et al.
Veröffentlicht: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
von: Lake, Thom, et al.
Veröffentlicht: (2024)
von: Lake, Thom, et al.
Veröffentlicht: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
A Long Way to Go: Investigating Length Correlations in RLHF
von: Singhal, Prasann, et al.
Veröffentlicht: (2023)
von: Singhal, Prasann, et al.
Veröffentlicht: (2023)
FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents
von: Pham, Quang Hieu, et al.
Veröffentlicht: (2026)
von: Pham, Quang Hieu, et al.
Veröffentlicht: (2026)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
von: Ye, Xi, et al.
Veröffentlicht: (2025)
von: Ye, Xi, et al.
Veröffentlicht: (2025)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
von: Wu, Rui, et al.
Veröffentlicht: (2026)
von: Wu, Rui, et al.
Veröffentlicht: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
von: Deng, Wenlong, et al.
Veröffentlicht: (2026)
von: Deng, Wenlong, et al.
Veröffentlicht: (2026)
Adaptive Margin RLHF via Preference over Preferences
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
von: Sriram, Aniruddh, et al.
Veröffentlicht: (2024)
von: Sriram, Aniruddh, et al.
Veröffentlicht: (2024)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
von: Bello, Femi, et al.
Veröffentlicht: (2025)
von: Bello, Femi, et al.
Veröffentlicht: (2025)
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing
von: Wei, Jiayi, et al.
Veröffentlicht: (2023)
von: Wei, Jiayi, et al.
Veröffentlicht: (2023)
A Gradient Analysis Framework for Rewarding Good and Penalizing Bad Examples in Language Models
von: Tuan, Yi-Lin, et al.
Veröffentlicht: (2024)
von: Tuan, Yi-Lin, et al.
Veröffentlicht: (2024)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
MIC: Maximizing Informational Capacity in Adaptive Representations via Isotropic Subspace Alignment
von: Hong, Dang Nguyen, et al.
Veröffentlicht: (2026)
von: Hong, Dang Nguyen, et al.
Veröffentlicht: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
High-Layer Attention Pruning with Rescaling
von: Liu, Songtao, et al.
Veröffentlicht: (2025)
von: Liu, Songtao, et al.
Veröffentlicht: (2025)
Reward-free Alignment for Conflicting Objectives
von: Chen, Peter, et al.
Veröffentlicht: (2026)
von: Chen, Peter, et al.
Veröffentlicht: (2026)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
von: Jang, Eyon, et al.
Veröffentlicht: (2026)
von: Jang, Eyon, et al.
Veröffentlicht: (2026)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
RewardAnything: Generalizable Principle-Following Reward Models
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
On Teacher Hacking in Language Model Distillation
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
von: Wang, Li, et al.
Veröffentlicht: (2026)
von: Wang, Li, et al.
Veröffentlicht: (2026)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
von: Bai, Yang, et al.
Veröffentlicht: (2026)
von: Bai, Yang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LoFiT: Localized Fine-tuning on LLM Representations
von: Yin, Fangcong, et al.
Veröffentlicht: (2024) -
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
von: Xin, Yutong, et al.
Veröffentlicht: (2026) -
Understanding Synthetic Context Extension via Retrieval Heads
von: Zhao, Xinyu, et al.
Veröffentlicht: (2024) -
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2024) -
CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation
von: Khatry, Anirudh, et al.
Veröffentlicht: (2025)