Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benavoli, Alessio, Facchini, Alessandro, Zaffalon, Marco
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918265380405248
author Benavoli, Alessio
Facchini, Alessandro
Zaffalon, Marco
author_facet Benavoli, Alessio
Facchini, Alessandro
Zaffalon, Marco
contents How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that helps a human to maximise their utility function(s). However, only the human knows these function(s); the AI assistant must learn them. The shutdown problem instead concerns designing AI agents that: shut down when a shutdown button is pressed; neither try to prevent nor cause the pressing of the shutdown button; and otherwise accomplish their task competently. In this paper, we show that addressing these challenges requires AI agents that can reason under uncertainty and handle both incomplete and non-Archimedean preferences.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23508
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities
Benavoli, Alessio
Facchini, Alessandro
Zaffalon, Marco
Artificial Intelligence
Computer Science and Game Theory
How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that helps a human to maximise their utility function(s). However, only the human knows these function(s); the AI assistant must learn them. The shutdown problem instead concerns designing AI agents that: shut down when a shutdown button is pressed; neither try to prevent nor cause the pressing of the shutdown button; and otherwise accomplish their task competently. In this paper, we show that addressing these challenges requires AI agents that can reason under uncertainty and handle both incomplete and non-Archimedean preferences.
title Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities
topic Artificial Intelligence
Computer Science and Game Theory
url https://arxiv.org/abs/2512.23508