Self-Attribution Bias: When AI Monitors Go Easy on Themselves

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Khullar, Dipika, Hopkins, Jack, Wang, Rowan, Roger, Fabien
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918371040165888
author Khullar, Dipika
Hopkins, Jack
Wang, Rowan
Roger, Fabien
author_facet Khullar, Dipika
Hopkins, Jack
Wang, Rowan
Roger, Fabien
contents Agentic systems increasingly rely on language models to monitor their own behavior. For example, coding agents may self critique generated code for pull request approval or assess the safety of tool-use actions. We show that this design pattern can fail when the action is presented in a previous or in the same assistant turn instead of being presented by the user in a user turn. We define self-attribution bias as the tendency of a model to evaluate an action as more correct or less risky when the action is implicitly framed as its own, compared to when the same action is evaluated under off-policy attribution. Across four coding and tool-use datasets, we find that monitors fail to report high-risk or low-correctness actions more often when evaluation follows a previous assistant turn in which the action was generated, compared to when the same action is evaluated in a new context presented in a user turn. In contrast, explicitly stating that the action comes from the monitor does not by itself induce self-attribution bias. Because monitors are often evaluated on fixed examples rather than on their own generated actions, these evaluations can make monitors appear more reliable than they actually are in deployment, leading developers to unknowingly deploy inadequate monitors in agentic systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04582
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Attribution Bias: When AI Monitors Go Easy on Themselves
Khullar, Dipika
Hopkins, Jack
Wang, Rowan
Roger, Fabien
Artificial Intelligence
Machine Learning
Agentic systems increasingly rely on language models to monitor their own behavior. For example, coding agents may self critique generated code for pull request approval or assess the safety of tool-use actions. We show that this design pattern can fail when the action is presented in a previous or in the same assistant turn instead of being presented by the user in a user turn. We define self-attribution bias as the tendency of a model to evaluate an action as more correct or less risky when the action is implicitly framed as its own, compared to when the same action is evaluated under off-policy attribution. Across four coding and tool-use datasets, we find that monitors fail to report high-risk or low-correctness actions more often when evaluation follows a previous assistant turn in which the action was generated, compared to when the same action is evaluated in a new context presented in a user turn. In contrast, explicitly stating that the action comes from the monitor does not by itself induce self-attribution bias. Because monitors are often evaluated on fixed examples rather than on their own generated actions, these evaluations can make monitors appear more reliable than they actually are in deployment, leading developers to unknowingly deploy inadequate monitors in agentic systems.
title Self-Attribution Bias: When AI Monitors Go Easy on Themselves
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.04582