Saved in:
Bibliographic Details
Main Authors: Ebtekar, Aram, Cohen, Michael K.
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.13609
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910130335907840
author Ebtekar, Aram
Cohen, Michael K.
author_facet Ebtekar, Aram
Cohen, Michael K.
contents Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward range to include a large negative value $-L$, while the true environment's rewards lie in $[0,1]$. After observing consistently high rewards, the Bayesian policy becomes risk-averse to novel schemes that plausibly lead to $-L$. We design a simple override mechanism that yields control to a safe mentor whenever the predicted value drops below a fixed threshold. We prove two properties of the resulting agent: (i) Capability: using mentor-guided exploration with vanishing frequency, the agent attains sublinear regret against its best mentor. (ii) Safety: no decidable low-complexity predicate is triggered by the optimizing policy before it is triggered by a mentor.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13609
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Golden Handcuffs make safer AI agents
Ebtekar, Aram
Cohen, Michael K.
Machine Learning
Artificial Intelligence
68Q32 (Primary) 68Q30 (Secondary)
I.2.6
Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward range to include a large negative value $-L$, while the true environment's rewards lie in $[0,1]$. After observing consistently high rewards, the Bayesian policy becomes risk-averse to novel schemes that plausibly lead to $-L$. We design a simple override mechanism that yields control to a safe mentor whenever the predicted value drops below a fixed threshold. We prove two properties of the resulting agent: (i) Capability: using mentor-guided exploration with vanishing frequency, the agent attains sublinear regret against its best mentor. (ii) Safety: no decidable low-complexity predicate is triggered by the optimizing policy before it is triggered by a mentor.
title Golden Handcuffs make safer AI agents
topic Machine Learning
Artificial Intelligence
68Q32 (Primary) 68Q30 (Secondary)
I.2.6
url https://arxiv.org/abs/2604.13609