Modeling Human Beliefs about AI Behavior for Scalable Oversight

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lang, Leon, Forré, Patrick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917029823381504
author Lang, Leon
Forré, Patrick
author_facet Lang, Leon
Forré, Patrick
contents As AI systems advance beyond human capabilities, scalable oversight becomes critical: how can we supervise AI that exceeds our abilities? A key challenge is that human evaluators may form incorrect beliefs about AI behavior in complex tasks, leading to unreliable feedback and poor value inference. To address this, we propose modeling evaluators' beliefs to interpret their feedback more reliably. We formalize human belief models, analyze their theoretical role in value learning, and characterize when ambiguity remains. To reduce reliance on precise belief models, we introduce "belief model covering" as a relaxation. This motivates our preliminary proposal to use the internal representations of adapted foundation models to mimic human evaluators' beliefs. These representations could be used to learn correct values from human feedback even when evaluators misunderstand the AI's behavior. Our work suggests that modeling human beliefs can improve value learning and outlines practical research directions for implementing this approach to scalable oversight.
format Preprint
id arxiv_https___arxiv_org_abs_2502_21262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modeling Human Beliefs about AI Behavior for Scalable Oversight
Lang, Leon
Forré, Patrick
Artificial Intelligence
Machine Learning
As AI systems advance beyond human capabilities, scalable oversight becomes critical: how can we supervise AI that exceeds our abilities? A key challenge is that human evaluators may form incorrect beliefs about AI behavior in complex tasks, leading to unreliable feedback and poor value inference. To address this, we propose modeling evaluators' beliefs to interpret their feedback more reliably. We formalize human belief models, analyze their theoretical role in value learning, and characterize when ambiguity remains. To reduce reliance on precise belief models, we introduce "belief model covering" as a relaxation. This motivates our preliminary proposal to use the internal representations of adapted foundation models to mimic human evaluators' beliefs. These representations could be used to learn correct values from human feedback even when evaluators misunderstand the AI's behavior. Our work suggests that modeling human beliefs can improve value learning and outlines practical research directions for implementing this approach to scalable oversight.
title Modeling Human Beliefs about AI Behavior for Scalable Oversight
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.21262