What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pan, Jinhao, Raj, Chahat, Yao, Ziyu, Zhu, Ziwei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909791474941952
author Pan, Jinhao
Raj, Chahat
Yao, Ziyu
Zhu, Ziwei
author_facet Pan, Jinhao
Raj, Chahat
Yao, Ziyu
Zhu, Ziwei
contents Large Language Models (LLMs) often exhibit social biases inherited from their training data. While existing benchmarks evaluate bias by term-based mode through direct term associations between demographic terms and bias terms, LLMs have become increasingly adept at avoiding biased responses, leading to seemingly low levels of bias. However, biases persist in subtler, contextually hidden forms that traditional benchmarks fail to capture. We introduce the Description-based Bias Benchmark (DBB), a novel dataset designed to assess bias at the semantic level that bias concepts are hidden within naturalistic, subtly framed contexts in real-world scenarios rather than superficial terms. We analyze six state-of-the-art LLMs, revealing that while models reduce bias in response at the term level, they continue to reinforce biases in nuanced settings. Data, code, and results are available at https://github.com/JP-25/Description-based-Bias-Benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs
Pan, Jinhao
Raj, Chahat
Yao, Ziyu
Zhu, Ziwei
Computation and Language
Large Language Models (LLMs) often exhibit social biases inherited from their training data. While existing benchmarks evaluate bias by term-based mode through direct term associations between demographic terms and bias terms, LLMs have become increasingly adept at avoiding biased responses, leading to seemingly low levels of bias. However, biases persist in subtler, contextually hidden forms that traditional benchmarks fail to capture. We introduce the Description-based Bias Benchmark (DBB), a novel dataset designed to assess bias at the semantic level that bias concepts are hidden within naturalistic, subtly framed contexts in real-world scenarios rather than superficial terms. We analyze six state-of-the-art LLMs, revealing that while models reduce bias in response at the term level, they continue to reinforce biases in nuanced settings. Data, code, and results are available at https://github.com/JP-25/Description-based-Bias-Benchmark.
title What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs
topic Computation and Language
url https://arxiv.org/abs/2502.19749