Shutdownable Agents through POST-Agency

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Thornley, Elliott
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915705470844928
author Thornley, Elliott
author_facet Thornley, Elliott
contents Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happen. I propose that we train agents to satisfy Preferences Only Between Same-Length Trajectories (POST). I then prove that POST - together with other conditions - implies Neutrality+: the agent maximizes expected utility, ignoring the probability distribution over trajectory-lengths. I argue that Neutrality+ keeps agents shutdownable and allows them to be useful.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20203
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Shutdownable Agents through POST-Agency
Thornley, Elliott
Artificial Intelligence
Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happen. I propose that we train agents to satisfy Preferences Only Between Same-Length Trajectories (POST). I then prove that POST - together with other conditions - implies Neutrality+: the agent maximizes expected utility, ignoring the probability distribution over trajectory-lengths. I argue that Neutrality+ keeps agents shutdownable and allows them to be useful.
title Shutdownable Agents through POST-Agency
topic Artificial Intelligence
url https://arxiv.org/abs/2505.20203