The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Konrad, Phongsakon Mark, Adam, Tim Lukas, Merrild, Ane Cathrine Holst, Terrenzi, Riccardo, De Rosa, Rebecca, Tanyel, Toygar, Ayvaz, Serkan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917480725741568
author Konrad, Phongsakon Mark
Adam, Tim Lukas
Merrild, Ane Cathrine Holst
Terrenzi, Riccardo
De Rosa, Rebecca
Tanyel, Toygar
Ayvaz, Serkan
author_facet Konrad, Phongsakon Mark
Adam, Tim Lukas
Merrild, Ane Cathrine Holst
Terrenzi, Riccardo
De Rosa, Rebecca
Tanyel, Toygar
Ayvaz, Serkan
contents AI deployment in sensitive domains such as health care, credit, employment, and criminal justice is often treated as unsafe to authorize until model internals can be explained. This often leads to an excessive reliance on mechanistic interpretability to address a deployment challenge beyond its intended scope. We argue that the gate should instead be calibrated verification: authorization should be domain-scoped, independently checkable, monitored after release, accountable, contestable, and revocable. The reason is twofold. First, model capability is uneven across nearby tasks, so authorization must attach to a specific use rather than to a model in general. Second, societies have long governed opaque expertise through credentials, monitoring, liability, appeal, and revocation rather than mechanism-level explanation. Recent evidence reinforces this distinction between mechanistic understanding and deployment authority: a 53-percentage-point gap between internal representations and output correction shows that understanding may not translate into action, while one scoping review found that only 9.0% of FDA-approved AI/ML device documents contained a prospective post-market surveillance study. We propose Verification Coverage, a six-component reportable standard with a minimum-composition rule, as the metric that should sit beside capability scores in model cards, leaderboards, and regulatory disclosures.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10601
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime
Konrad, Phongsakon Mark
Adam, Tim Lukas
Merrild, Ane Cathrine Holst
Terrenzi, Riccardo
De Rosa, Rebecca
Tanyel, Toygar
Ayvaz, Serkan
Artificial Intelligence
AI deployment in sensitive domains such as health care, credit, employment, and criminal justice is often treated as unsafe to authorize until model internals can be explained. This often leads to an excessive reliance on mechanistic interpretability to address a deployment challenge beyond its intended scope. We argue that the gate should instead be calibrated verification: authorization should be domain-scoped, independently checkable, monitored after release, accountable, contestable, and revocable. The reason is twofold. First, model capability is uneven across nearby tasks, so authorization must attach to a specific use rather than to a model in general. Second, societies have long governed opaque expertise through credentials, monitoring, liability, appeal, and revocation rather than mechanism-level explanation. Recent evidence reinforces this distinction between mechanistic understanding and deployment authority: a 53-percentage-point gap between internal representations and output correction shows that understanding may not translate into action, while one scoping review found that only 9.0% of FDA-approved AI/ML device documents contained a prospective post-market surveillance study. We propose Verification Coverage, a six-component reportable standard with a minimum-composition rule, as the metric that should sit beside capability scores in model cards, leaderboards, and regulatory disclosures.
title The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime
topic Artificial Intelligence
url https://arxiv.org/abs/2605.10601