Off-Switching Not Guaranteed
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913689150423040 |
|---|---|
| author | Neth, Sven |
| author_facet | Neth, Sven |
| contents | Hadfield-Menell et al. (2017) propose the Off-Switch Game, a model of Human-AI cooperation in which AI agents always defer to humans because they are uncertain about our preferences. I explain two reasons why AI agents might not defer. First, AI agents might not value learning. Second, even if AI agents value learning, they might not be certain to learn our actual preferences. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_08864 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Off-Switching Not Guaranteed Neth, Sven Artificial Intelligence Hadfield-Menell et al. (2017) propose the Off-Switch Game, a model of Human-AI cooperation in which AI agents always defer to humans because they are uncertain about our preferences. I explain two reasons why AI agents might not defer. First, AI agents might not value learning. Second, even if AI agents value learning, they might not be certain to learn our actual preferences. |
| title | Off-Switching Not Guaranteed |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2502.08864 |