Inference-Time Reward Hacking in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Khalaf, Hadi, Verdun, Claudio Mayrink, Oesterling, Alex, Lakkaraju, Himabindu, Calmon, Flavio du Pin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
Soft Best-of-n Sampling for Model Alignment
by: Verdun, Claudio Mayrink, et al.
Published: (2025)
by: Verdun, Claudio Mayrink, et al.
Published: (2025)
AI Alignment at Your Discretion
by: Buyl, Maarten, et al.
Published: (2025)
by: Buyl, Maarten, et al.
Published: (2025)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Fair Machine Unlearning: Data Removal while Mitigating Disparities
by: Oesterling, Alex, et al.
Published: (2023)
by: Oesterling, Alex, et al.
Published: (2023)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
by: Badrinath, Charumathi, et al.
Published: (2024)
by: Badrinath, Charumathi, et al.
Published: (2024)
Multi-Group Proportional Representation in Retrieval
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
Multi-Group Proportional Representation for Text-to-Image Models
by: Jung, Sangwon, et al.
Published: (2025)
by: Jung, Sangwon, et al.
Published: (2025)
Robust AI Evaluation through Maximal Lotteries
by: Khalaf, Hadi, et al.
Published: (2026)
by: Khalaf, Hadi, et al.
Published: (2026)
Explaining the Model, Protecting Your Data: Revealing and Mitigating the Data Privacy Risks of Post-Hoc Model Explanations via Membership Inference
by: Huang, Catherine, et al.
Published: (2024)
by: Huang, Catherine, et al.
Published: (2024)
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications
by: Liu, Yanchen, et al.
Published: (2023)
by: Liu, Yanchen, et al.
Published: (2023)
In-Context Unlearning: Language Models as Few Shot Unlearners
by: Pawelczyk, Martin, et al.
Published: (2023)
by: Pawelczyk, Martin, et al.
Published: (2023)
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024)
by: Rawal, Kaivalya, et al.
Published: (2024)
Optimized Couplings for Watermarking Large Language Models
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
Characterizing Data Point Vulnerability via Average-Case Robustness
by: Han, Tessa, et al.
Published: (2023)
by: Han, Tessa, et al.
Published: (2023)
On the Trade-offs between Adversarial Robustness and Actionable Explanations
by: Krishna, Satyapriya, et al.
Published: (2023)
by: Krishna, Satyapriya, et al.
Published: (2023)
GradPCA: Leveraging NTK Alignment for Reliable Out-of-Distribution Detection
by: Seleznova, Mariia, et al.
Published: (2025)
by: Seleznova, Mariia, et al.
Published: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
by: Xiong, Zidi, et al.
Published: (2026)
by: Xiong, Zidi, et al.
Published: (2026)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
Predictive Churn with the Set of Good Models
by: Watson-Daniels, Jamelle, et al.
Published: (2024)
by: Watson-Daniels, Jamelle, et al.
Published: (2024)
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
by: Bhalla, Usha, et al.
Published: (2023)
by: Bhalla, Usha, et al.
Published: (2023)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
by: Zhai, Kevin, et al.
Published: (2025)
by: Zhai, Kevin, et al.
Published: (2025)
Non-Asymptotic Uncertainty Quantification in High-Dimensional Learning
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
Towards Unifying Interpretability and Control: Evaluation via Intervention
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Attack-Aware Noise Calibration for Differential Privacy
by: Kulynych, Bogdan, et al.
Published: (2024)
by: Kulynych, Bogdan, et al.
Published: (2024)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Inference-Time Machine Unlearning via Gated Activation Redirection
by: Turani, Vinícius Conte, et al.
Published: (2026)
by: Turani, Vinícius Conte, et al.
Published: (2026)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023)
by: Eisenstein, Jacob, et al.
Published: (2023)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
High-Dimensional Confidence Regions in Sparse MRI
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
With or Without Replacement? Improving Confidence in Fourier Imaging
by: Hoppe, Frederik, et al.
Published: (2024)
by: Hoppe, Frederik, et al.
Published: (2024)
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
by: Krishna, Satyapriya, et al.
Published: (2022)
by: Krishna, Satyapriya, et al.
Published: (2022)
Defining and Characterizing Reward Hacking
by: Skalse, Joar, et al.
Published: (2022)
by: Skalse, Joar, et al.
Published: (2022)
Manipulating Large Language Models to Increase Product Visibility
by: Kumar, Aounon, et al.
Published: (2024)
by: Kumar, Aounon, et al.
Published: (2024)
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Similar Items
-
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025) -
Soft Best-of-n Sampling for Model Alignment
by: Verdun, Claudio Mayrink, et al.
Published: (2025) -
AI Alignment at Your Discretion
by: Buyl, Maarten, et al.
Published: (2025) -
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
by: Bhalla, Usha, et al.
Published: (2024) -
Fair Machine Unlearning: Data Removal while Mitigating Disparities
by: Oesterling, Alex, et al.
Published: (2023)