Textual Entailment is not a Better Bias Metric than Token Probability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Felkner, Virginia K., Lim, Allison, May, Jonathan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914255893168128
author Felkner, Virginia K.
Lim, Allison
May, Jonathan
author_facet Felkner, Virginia K.
Lim, Allison
May, Jonathan
contents Measurement of social bias in language models is typically by token probability (TP) metrics, which are broadly applicable but have been criticized for their distance from real-world language model use cases and harms. In this work, we test natural language inference (NLI) as an alternative bias metric. In extensive experiments across seven LM families, we show that NLI and TP bias evaluation behave substantially differently, with very low correlation among different NLI metrics and between NLI and TP metrics. NLI metrics are more brittle and unstable, slightly less sensitive to wording of counterstereotypical sentences, and slightly more sensitive to wording of tested stereotypes than TP approaches. Given this conflicting evidence, we conclude that neither token probability nor natural language inference is a ``better'' bias metric in all cases. We do not find sufficient evidence to justify NLI as a complete replacement for TP metrics in bias evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07662
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Textual Entailment is not a Better Bias Metric than Token Probability
Felkner, Virginia K.
Lim, Allison
May, Jonathan
Computation and Language
Computers and Society
I.2.7; K.4.2
Measurement of social bias in language models is typically by token probability (TP) metrics, which are broadly applicable but have been criticized for their distance from real-world language model use cases and harms. In this work, we test natural language inference (NLI) as an alternative bias metric. In extensive experiments across seven LM families, we show that NLI and TP bias evaluation behave substantially differently, with very low correlation among different NLI metrics and between NLI and TP metrics. NLI metrics are more brittle and unstable, slightly less sensitive to wording of counterstereotypical sentences, and slightly more sensitive to wording of tested stereotypes than TP approaches. Given this conflicting evidence, we conclude that neither token probability nor natural language inference is a ``better'' bias metric in all cases. We do not find sufficient evidence to justify NLI as a complete replacement for TP metrics in bias evaluation.
title Textual Entailment is not a Better Bias Metric than Token Probability
topic Computation and Language
Computers and Society
I.2.7; K.4.2
url https://arxiv.org/abs/2510.07662