Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Deng, Wenlong, Huang, Jiaji, Ozkara, Kaan, Li, Yushu, Thrampoulidis, Christos, Li, Xiaoxiao, Park, Youngsuk |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
por: Deng, Wenlong, et al.
Publicado: (2025)
por: Deng, Wenlong, et al.
Publicado: (2025)
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
por: Deng, Wenlong, et al.
Publicado: (2025)
por: Deng, Wenlong, et al.
Publicado: (2025)
Stochastic Rounding for LLM Training: Theory and Practice
por: Ozkara, Kaan, et al.
Publicado: (2025)
por: Ozkara, Kaan, et al.
Publicado: (2025)
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
por: Deng, Wenlong, et al.
Publicado: (2023)
por: Deng, Wenlong, et al.
Publicado: (2023)
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
por: Deng, Wenlong, et al.
Publicado: (2024)
por: Deng, Wenlong, et al.
Publicado: (2024)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
por: Thrampoulidis, Christos
Publicado: (2024)
por: Thrampoulidis, Christos
Publicado: (2024)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
por: Deng, Wenlong, et al.
Publicado: (2024)
por: Deng, Wenlong, et al.
Publicado: (2024)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
por: Thrampoulidis, Christos, et al.
Publicado: (2025)
por: Thrampoulidis, Christos, et al.
Publicado: (2025)
On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
por: Deng, Wenlong, et al.
Publicado: (2025)
por: Deng, Wenlong, et al.
Publicado: (2025)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
por: Khaled, Ahmed, et al.
Publicado: (2025)
por: Khaled, Ahmed, et al.
Publicado: (2025)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
por: Behnia, Tina, et al.
Publicado: (2025)
por: Behnia, Tina, et al.
Publicado: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
por: Zhao, Yize, et al.
Publicado: (2024)
por: Zhao, Yize, et al.
Publicado: (2024)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
por: Liu, Hongyi, et al.
Publicado: (2025)
por: Liu, Hongyi, et al.
Publicado: (2025)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
por: Wu, Rui, et al.
Publicado: (2026)
por: Wu, Rui, et al.
Publicado: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
por: Li, Yushu, et al.
Publicado: (2026)
por: Li, Yushu, et al.
Publicado: (2026)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
por: Wang, Songtao, et al.
Publicado: (2026)
por: Wang, Songtao, et al.
Publicado: (2026)
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
por: Lu, Keming, et al.
Publicado: (2024)
por: Lu, Keming, et al.
Publicado: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
SPIRE: Conditional Personalization for Federated Diffusion Generative Models
por: Ozkara, Kaan, et al.
Publicado: (2025)
por: Ozkara, Kaan, et al.
Publicado: (2025)
ADEPT: Hierarchical Bayes Approach to Personalized Federated Unsupervised Learning
por: Ozkara, Kaan, et al.
Publicado: (2024)
por: Ozkara, Kaan, et al.
Publicado: (2024)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
por: Zhai, Kevin, et al.
Publicado: (2025)
por: Zhai, Kevin, et al.
Publicado: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
por: Wang, Chaoqi, et al.
Publicado: (2025)
por: Wang, Chaoqi, et al.
Publicado: (2025)
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
por: Su, Zeli, et al.
Publicado: (2026)
por: Su, Zeli, et al.
Publicado: (2026)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
por: Pang, Jing-Cheng, et al.
Publicado: (2024)
por: Pang, Jing-Cheng, et al.
Publicado: (2024)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
por: Deora, Puneesh, et al.
Publicado: (2025)
por: Deora, Puneesh, et al.
Publicado: (2025)
Supervised Contrastive Representation Learning: Landscape Analysis with Unconstrained Features
por: Behnia, Tina, et al.
Publicado: (2024)
por: Behnia, Tina, et al.
Publicado: (2024)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
por: Vasudeva, Bhavya, et al.
Publicado: (2026)
por: Vasudeva, Bhavya, et al.
Publicado: (2026)
Transformers as Support Vector Machines
por: Tarzanagh, Davoud Ataee, et al.
Publicado: (2023)
por: Tarzanagh, Davoud Ataee, et al.
Publicado: (2023)
Fine-Tuning Language Models with Reward Learning on Policy
por: Lang, Hao, et al.
Publicado: (2024)
por: Lang, Hao, et al.
Publicado: (2024)
Noise Contrastive Alignment of Language Models with Explicit Rewards
por: Chen, Huayu, et al.
Publicado: (2024)
por: Chen, Huayu, et al.
Publicado: (2024)
Energy-Based Reward Models for Robust Language Model Alignment
por: Lochab, Anamika, et al.
Publicado: (2025)
por: Lochab, Anamika, et al.
Publicado: (2025)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
por: Liu, Hongyi, et al.
Publicado: (2025)
por: Liu, Hongyi, et al.
Publicado: (2025)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
Ejemplares similares
-
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
por: Deng, Wenlong, et al.
Publicado: (2025) -
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
por: Deng, Wenlong, et al.
Publicado: (2025) -
Stochastic Rounding for LLM Training: Theory and Practice
por: Ozkara, Kaan, et al.
Publicado: (2025) -
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
por: Deng, Wenlong, et al.
Publicado: (2023) -
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
por: Deng, Wenlong, et al.
Publicado: (2024)