Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.13609 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910130335907840 |
|---|---|
| author | Ebtekar, Aram Cohen, Michael K. |
| author_facet | Ebtekar, Aram Cohen, Michael K. |
| contents | Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward range to include a large negative value $-L$, while the true environment's rewards lie in $[0,1]$. After observing consistently high rewards, the Bayesian policy becomes risk-averse to novel schemes that plausibly lead to $-L$. We design a simple override mechanism that yields control to a safe mentor whenever the predicted value drops below a fixed threshold. We prove two properties of the resulting agent: (i) Capability: using mentor-guided exploration with vanishing frequency, the agent attains sublinear regret against its best mentor. (ii) Safety: no decidable low-complexity predicate is triggered by the optimizing policy before it is triggered by a mentor. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_13609 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Golden Handcuffs make safer AI agents Ebtekar, Aram Cohen, Michael K. Machine Learning Artificial Intelligence 68Q32 (Primary) 68Q30 (Secondary) I.2.6 Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward range to include a large negative value $-L$, while the true environment's rewards lie in $[0,1]$. After observing consistently high rewards, the Bayesian policy becomes risk-averse to novel schemes that plausibly lead to $-L$. We design a simple override mechanism that yields control to a safe mentor whenever the predicted value drops below a fixed threshold. We prove two properties of the resulting agent: (i) Capability: using mentor-guided exploration with vanishing frequency, the agent attains sublinear regret against its best mentor. (ii) Safety: no decidable low-complexity predicate is triggered by the optimizing policy before it is triggered by a mentor. |
| title | Golden Handcuffs make safer AI agents |
| topic | Machine Learning Artificial Intelligence 68Q32 (Primary) 68Q30 (Secondary) I.2.6 |
| url | https://arxiv.org/abs/2604.13609 |