Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Deng, Wenlong, Huang, Jiaji, Ozkara, Kaan, Li, Yushu, Thrampoulidis, Christos, Li, Xiaoxiao, Park, Youngsuk
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916043478269952
author Deng, Wenlong
Huang, Jiaji
Ozkara, Kaan
Li, Yushu
Thrampoulidis, Christos
Li, Xiaoxiao
Park, Youngsuk
author_facet Deng, Wenlong
Huang, Jiaji
Ozkara, Kaan
Li, Yushu
Thrampoulidis, Christos
Li, Xiaoxiao
Park, Youngsuk
contents Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode through the geometry of reinforcement learning updates in language models and argue that hacking emerges when optimization drifts away from a stable low-dimensional learning trajectory. We analyze this drift through dominant singular directions of parameter updates and show that reward-hacking runs exhibit substantially larger directional change than clean runs. Motivated by this observation, we introduce trusted-direction projection, which constrains gradients to remain within a clean reference subspace. Across reward-hacking experiments on mathematical reasoning, the proposed approach delays shortcut exploitation and better preserves task performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25189
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
Deng, Wenlong
Huang, Jiaji
Ozkara, Kaan
Li, Yushu
Thrampoulidis, Christos
Li, Xiaoxiao
Park, Youngsuk
Machine Learning
Computation and Language
Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode through the geometry of reinforcement learning updates in language models and argue that hacking emerges when optimization drifts away from a stable low-dimensional learning trajectory. We analyze this drift through dominant singular directions of parameter updates and show that reward-hacking runs exhibit substantially larger directional change than clean runs. Motivated by this observation, we introduce trusted-direction projection, which constrains gradients to remain within a clean reference subspace. Across reward-hacking experiments on mathematical reasoning, the proposed approach delays shortcut exploitation and better preserves task performance.
title Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.25189