Reward Hacking Mitigation using Verifiable Composite Rewards

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tarek, Mirza Farhan Bin, Beheshti, Rahmatollah
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908546760704000
author Tarek, Mirza Farhan Bin
Beheshti, Rahmatollah
author_facet Tarek, Mirza Farhan Bin
Beheshti, Rahmatollah
contents Reinforcement Learning from Verifiable Rewards (RLVR) has recently shown that large language models (LLMs) can develop their own reasoning without direct supervision. However, applications in the medical domain, specifically for question answering, are susceptible to significant reward hacking during the reasoning phase. Our work addresses two primary forms of this behavior: i) providing a final answer without preceding reasoning, and ii) employing non-standard reasoning formats to exploit the reward mechanism. To mitigate these, we introduce a composite reward function with specific penalties for these behaviors. Our experiments show that extending RLVR with our proposed reward model leads to better-formatted reasoning with less reward hacking and good accuracy compared to the baselines. This approach marks a step toward reducing reward hacking and enhancing the reliability of models utilizing RLVR.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15557
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reward Hacking Mitigation using Verifiable Composite Rewards
Tarek, Mirza Farhan Bin
Beheshti, Rahmatollah
Machine Learning
Artificial Intelligence
Reinforcement Learning from Verifiable Rewards (RLVR) has recently shown that large language models (LLMs) can develop their own reasoning without direct supervision. However, applications in the medical domain, specifically for question answering, are susceptible to significant reward hacking during the reasoning phase. Our work addresses two primary forms of this behavior: i) providing a final answer without preceding reasoning, and ii) employing non-standard reasoning formats to exploit the reward mechanism. To mitigate these, we introduce a composite reward function with specific penalties for these behaviors. Our experiments show that extending RLVR with our proposed reward model leads to better-formatted reasoning with less reward hacking and good accuracy compared to the baselines. This approach marks a step toward reducing reward hacking and enhancing the reliability of models utilizing RLVR.
title Reward Hacking Mitigation using Verifiable Composite Rewards
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.15557