PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jing-Jing, Mire, Joel, Fleisig, Eve, Pyatkin, Valentina, Collins, Anne, Sap, Maarten, Levine, Sydney
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914304362545152
author Li, Jing-Jing
Mire, Joel
Fleisig, Eve
Pyatkin, Valentina
Collins, Anne
Sap, Maarten
Levine, Sydney
author_facet Li, Jing-Jing
Mire, Joel
Fleisig, Eve
Pyatkin, Valentina
Collins, Anne
Sap, Maarten
Levine, Sydney
contents Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduce PluriHarms, a benchmark designed to systematically study human harm judgments across two key dimensions -- the harm axis (benign to harmful) and the agreement axis (agreement to disagreement). Our scalable framework generates prompts that capture diverse AI harms and human values while targeting cases with high disagreement rates, validated by human data. The benchmark includes 150 prompts with 15,000 ratings from 100 human annotators, enriched with demographic and psychological traits and prompt-level features of harmful actions, effects, and values. Our analyses show that prompts that relate to imminent risks and tangible harms amplify perceived harmfulness, while annotator traits (e.g., toxicity experience, education) and their interactions with prompt content explain systematic disagreement. We benchmark AI safety models and alignment methods on PluriHarms, finding that while personalization significantly improves prediction of human harm judgments, considerable room remains for future progress. By explicitly targeting value diversity and disagreement, our work provides a principled benchmark for moving beyond "one-size-fits-all" safety toward pluralistically safe AI.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08951
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
Li, Jing-Jing
Mire, Joel
Fleisig, Eve
Pyatkin, Valentina
Collins, Anne
Sap, Maarten
Levine, Sydney
Computers and Society
Artificial Intelligence
Computation and Language
Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduce PluriHarms, a benchmark designed to systematically study human harm judgments across two key dimensions -- the harm axis (benign to harmful) and the agreement axis (agreement to disagreement). Our scalable framework generates prompts that capture diverse AI harms and human values while targeting cases with high disagreement rates, validated by human data. The benchmark includes 150 prompts with 15,000 ratings from 100 human annotators, enriched with demographic and psychological traits and prompt-level features of harmful actions, effects, and values. Our analyses show that prompts that relate to imminent risks and tangible harms amplify perceived harmfulness, while annotator traits (e.g., toxicity experience, education) and their interactions with prompt content explain systematic disagreement. We benchmark AI safety models and alignment methods on PluriHarms, finding that while personalization significantly improves prediction of human harm judgments, considerable room remains for future progress. By explicitly targeting value diversity and disagreement, our work provides a principled benchmark for moving beyond "one-size-fits-all" safety toward pluralistically safe AI.
title PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
topic Computers and Society
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.08951