Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vaccaro, Michelle, Song, Jaeyoon, Almaatouq, Abdullah, Bakker, Michiel A.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918413451919360
author Vaccaro, Michelle
Song, Jaeyoon
Almaatouq, Abdullah
Bakker, Michiel A.
author_facet Vaccaro, Michelle
Song, Jaeyoon
Almaatouq, Abdullah
Bakker, Michiel A.
contents Current frontier AI safety evaluations emphasize static benchmarks, third-party annotations, and red-teaming. In this position paper, we argue that AI safety research should focus on human-centered evaluations that measure harmful capability uplift: the marginal increase in a user's ability to cause harm with a frontier model beyond what conventional tools already enable. We frame harmful capability uplift as a core AI safety metric, ground it in prior social science research, and provide concrete methodological guidance for systematic measurement. We conclude with actionable steps for developers, researchers, funders, and regulators to make harmful capability uplift evaluation a standard practice.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26676
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
Vaccaro, Michelle
Song, Jaeyoon
Almaatouq, Abdullah
Bakker, Michiel A.
Computers and Society
Artificial Intelligence
Human-Computer Interaction
Current frontier AI safety evaluations emphasize static benchmarks, third-party annotations, and red-teaming. In this position paper, we argue that AI safety research should focus on human-centered evaluations that measure harmful capability uplift: the marginal increase in a user's ability to cause harm with a frontier model beyond what conventional tools already enable. We frame harmful capability uplift as a core AI safety metric, ground it in prior social science research, and provide concrete methodological guidance for systematic measurement. We conclude with actionable steps for developers, researchers, funders, and regulators to make harmful capability uplift evaluation a standard practice.
title Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
topic Computers and Society
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2603.26676