MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Farquhar, Sebastian, Varma, Vikrant, Lindner, David, Elson, David, Biddulph, Caleb, Goodfellow, Ian, Shah, Rohin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910908498837504
author Farquhar, Sebastian
Varma, Vikrant
Lindner, David
Elson, David
Biddulph, Caleb
Goodfellow, Ian
Shah, Rohin
author_facet Farquhar, Sebastian
Varma, Vikrant
Lindner, David
Elson, David
Biddulph, Caleb
Goodfellow, Ian
Shah, Rohin
contents Future advanced AI systems may learn sophisticated strategies through reinforcement learning (RL) that humans cannot understand well enough to safely evaluate. We propose a training method which avoids agents learning undesired multi-step plans that receive high reward (multi-step "reward hacks") even if humans are not able to detect that the behaviour is undesired. The method, Myopic Optimization with Non-myopic Approval (MONA), works by combining short-sighted optimization with far-sighted reward. We demonstrate that MONA can prevent multi-step reward hacking that ordinary RL causes, even without being able to detect the reward hacking and without any extra information that ordinary RL does not get access to. We study MONA empirically in three settings which model different misalignment failure modes including 2-step environments with LLMs representing delegated oversight and encoded reasoning and longer-horizon gridworld environments representing sensor tampering.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13011
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
Farquhar, Sebastian
Varma, Vikrant
Lindner, David
Elson, David
Biddulph, Caleb
Goodfellow, Ian
Shah, Rohin
Machine Learning
Artificial Intelligence
Future advanced AI systems may learn sophisticated strategies through reinforcement learning (RL) that humans cannot understand well enough to safely evaluate. We propose a training method which avoids agents learning undesired multi-step plans that receive high reward (multi-step "reward hacks") even if humans are not able to detect that the behaviour is undesired. The method, Myopic Optimization with Non-myopic Approval (MONA), works by combining short-sighted optimization with far-sighted reward. We demonstrate that MONA can prevent multi-step reward hacking that ordinary RL causes, even without being able to detect the reward hacking and without any extra information that ordinary RL does not get access to. We study MONA empirically in three settings which model different misalignment failure modes including 2-step environments with LLMs representing delegated oversight and encoded reasoning and longer-horizon gridworld environments representing sensor tampering.
title MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2501.13011