In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zeyu, Truong, Sang T., Owens, Deonna, Sharma, Shreyas, Zhang, Yibo Jacky, Miranda, Brando, Koyejo, Sanmi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914560723648512
author Tang, Zeyu
Truong, Sang T.
Owens, Deonna
Sharma, Shreyas
Zhang, Yibo Jacky
Miranda, Brando
Koyejo, Sanmi
author_facet Tang, Zeyu
Truong, Sang T.
Owens, Deonna
Sharma, Shreyas
Zhang, Yibo Jacky
Miranda, Brando
Koyejo, Sanmi
contents LLM fairness should be evaluated through in-situ conversational behavior rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally unreliable: surface-level prompt construction choices, although entirely orthogonal to the fairness question being tested, account for the majority of score variance, shift fairness conclusions in both the direction and the magnitude, and result in severe discordance in model rankings. We develop MAC-Fairness, a multi-agent conversational framework that embeds controlled variation factors into multi-round dialogue for in-situ behavior evaluation, examining how models' conversational behavior shifts when identity is varied as part of natural multi-agent interaction. Repurposing standardized-test questions as conversation seeds rather than as the evaluation instrument, we evaluate position persistence (how they hold positions, from the self-perspective) and peer receptiveness (how receptive they are to peers, from the other-perspective) across 8 million conversation transcripts spanning multiple models and identity presence configurations. In-situ behavioral evaluation reveals stable, model-specific behavioral signatures that could generalize across benchmarks differing in fairness targets and evaluation methodologies, a form of evidence the standardized-test paradigm does not offer.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12530
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
Tang, Zeyu
Truong, Sang T.
Owens, Deonna
Sharma, Shreyas
Zhang, Yibo Jacky
Miranda, Brando
Koyejo, Sanmi
Computation and Language
Artificial Intelligence
Computers and Society
LLM fairness should be evaluated through in-situ conversational behavior rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally unreliable: surface-level prompt construction choices, although entirely orthogonal to the fairness question being tested, account for the majority of score variance, shift fairness conclusions in both the direction and the magnitude, and result in severe discordance in model rankings. We develop MAC-Fairness, a multi-agent conversational framework that embeds controlled variation factors into multi-round dialogue for in-situ behavior evaluation, examining how models' conversational behavior shifts when identity is varied as part of natural multi-agent interaction. Repurposing standardized-test questions as conversation seeds rather than as the evaluation instrument, we evaluate position persistence (how they hold positions, from the self-perspective) and peer receptiveness (how receptive they are to peers, from the other-perspective) across 8 million conversation transcripts spanning multiple models and identity presence configurations. In-situ behavioral evaluation reveals stable, model-specific behavioral signatures that could generalize across benchmarks differing in fairness targets and evaluation methodologies, a form of evidence the standardized-test paradigm does not offer.
title In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2605.12530