Friend or Foe: Delegating to an AI Whose Alignment is Unknown

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fudenberg, Drew, Liang, Annie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915500527714304
author Fudenberg, Drew
Liang, Annie
author_facet Fudenberg, Drew
Liang, Annie
contents AI systems have the potential to improve decision-making, but decision makers face the risk that the AI may be misaligned with their objectives. We study this problem in the context of a treatment decision, where a designer decides which patient attributes to reveal to an AI before receiving a prediction of the patient's need for treatment. Providing the AI with more information increases the benefits of an aligned AI but also amplifies the harm from a misaligned one. We characterize how the designer should select attributes to balance these competing forces, depending on their beliefs about the AI's reliability. We show that the designer should optimally disclose attributes that identify \emph{rare} segments of the population in which the need for treatment is high, and pool the remaining patients.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14396
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Friend or Foe: Delegating to an AI Whose Alignment is Unknown
Fudenberg, Drew
Liang, Annie
Theoretical Economics
Computer Science and Game Theory
AI systems have the potential to improve decision-making, but decision makers face the risk that the AI may be misaligned with their objectives. We study this problem in the context of a treatment decision, where a designer decides which patient attributes to reveal to an AI before receiving a prediction of the patient's need for treatment. Providing the AI with more information increases the benefits of an aligned AI but also amplifies the harm from a misaligned one. We characterize how the designer should select attributes to balance these competing forces, depending on their beliefs about the AI's reliability. We show that the designer should optimally disclose attributes that identify \emph{rare} segments of the population in which the need for treatment is high, and pool the remaining patients.
title Friend or Foe: Delegating to an AI Whose Alignment is Unknown
topic Theoretical Economics
Computer Science and Game Theory
url https://arxiv.org/abs/2509.14396