Do Large Language Models Reflect Demographic Pluralism in Safety?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Naseem, Usman, Kashyap, Gautam Siddharth, Ray, Sushant Kumar, Ali, Rafiq, Shabbir, Ebad, Mohammad, Abdullah
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908819083231232
author Naseem, Usman
Kashyap, Gautam Siddharth
Ray, Sushant Kumar
Ali, Rafiq
Shabbir, Ebad
Mohammad, Abdullah
author_facet Naseem, Usman
Kashyap, Gautam Siddharth
Ray, Sushant Kumar
Ali, Rafiq
Shabbir, Ebad
Mohammad, Abdullah
contents Large Language Model (LLM) safety is inherently pluralistic, reflecting variations in moral norms, cultural expectations, and demographic contexts. Yet, existing alignment datasets such as ANTHROPIC-HH and DICES rely on demographically narrow annotator pools, overlooking variation in safety perception across communities. Demo-SafetyBench addresses this gap by modeling demographic pluralism directly at the prompt level, decoupling value framing from responses. In Stage I, prompts from DICES are reclassified into 14 safety domains (adapted from BEAVERTAILS) using Mistral 7B-Instruct-v0.3, retaining demographic metadata and expanding low-resource domains via Llama-3.1-8B-Instruct with SimHash-based deduplication, yielding 43,050 samples. In Stage II, pluralistic sensitivity is evaluated using LLMs-as-Raters-Gemma-7B, GPT-4o, and LLaMA-2-7B-under zero-shot inference. Balanced thresholds (delta = 0.5, tau = 10) achieve high reliability (ICC = 0.87) and low demographic sensitivity (DS = 0.12), confirming that pluralistic safety evaluation can be both scalable and demographically robust.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07376
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do Large Language Models Reflect Demographic Pluralism in Safety?
Naseem, Usman
Kashyap, Gautam Siddharth
Ray, Sushant Kumar
Ali, Rafiq
Shabbir, Ebad
Mohammad, Abdullah
Computation and Language
Large Language Model (LLM) safety is inherently pluralistic, reflecting variations in moral norms, cultural expectations, and demographic contexts. Yet, existing alignment datasets such as ANTHROPIC-HH and DICES rely on demographically narrow annotator pools, overlooking variation in safety perception across communities. Demo-SafetyBench addresses this gap by modeling demographic pluralism directly at the prompt level, decoupling value framing from responses. In Stage I, prompts from DICES are reclassified into 14 safety domains (adapted from BEAVERTAILS) using Mistral 7B-Instruct-v0.3, retaining demographic metadata and expanding low-resource domains via Llama-3.1-8B-Instruct with SimHash-based deduplication, yielding 43,050 samples. In Stage II, pluralistic sensitivity is evaluated using LLMs-as-Raters-Gemma-7B, GPT-4o, and LLaMA-2-7B-under zero-shot inference. Balanced thresholds (delta = 0.5, tau = 10) achieve high reliability (ICC = 0.87) and low demographic sensitivity (DS = 0.12), confirming that pluralistic safety evaluation can be both scalable and demographically robust.
title Do Large Language Models Reflect Demographic Pluralism in Safety?
topic Computation and Language
url https://arxiv.org/abs/2602.07376