All Public Voices Are Equal, But Are Some More Equal Than Others to LLMs?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Sola, Janssen, Marco A., Wang, Jieshu, Min-Venditti, Ame, Karanjia, Neha, Anderies, John M.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910145780383744
author Kim, Sola
Janssen, Marco A.
Wang, Jieshu
Min-Venditti, Ame
Karanjia, Neha
Anderies, John M.
author_facet Kim, Sola
Janssen, Marco A.
Wang, Jieshu
Min-Venditti, Ame
Karanjia, Neha
Anderies, John M.
contents Federal agencies are increasingly deploying large language models (LLMs) to process public comments submitted during notice-and-comment rulemaking, the primary mechanism through which citizens influence federal regulation. Whether these systems treat all public input equally remains largely untested. Using a counterfactual design, we held comment content constant and varied only the commenter's demographic attribution -- race, gender, and socioeconomic status -- to test whether eight LLMs available for federal use produce differential summaries of identical comments. We processed 182 public comments across 32 identity conditions, generating over 106,000 summaries. Occupation was the only identity signal to produce consistent differential treatment: the same comment attributed to a street vendor, compared to a financial analyst, received a summary that preserved less of the original meaning, used simpler language, and shifted emotional tone. This pattern held across all names, prompts, models, and regulatory contexts tested. Race effects were inconsistent and appeared driven by specific name tokens rather than racial categories; gender effects were absent. Writing quality predicted summarization outcomes through argument substance rather than surface mechanics; experimentally injected spelling and grammar errors had negligible effects. The magnitude of occupation-based differential treatment varied by model provider, meaning that selecting a model implicitly selects a level of fairness -- a dimension that current procurement frameworks such as FedRAMP do not evaluate. These findings suggest that socioeconomic signals warrant attention in AI fairness assessments for government information systems, and that fairness benchmarks could be incorporated into existing federal IT procurement processes.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17247
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle All Public Voices Are Equal, But Are Some More Equal Than Others to LLMs?
Kim, Sola
Janssen, Marco A.
Wang, Jieshu
Min-Venditti, Ame
Karanjia, Neha
Anderies, John M.
Computers and Society
Human-Computer Interaction
Federal agencies are increasingly deploying large language models (LLMs) to process public comments submitted during notice-and-comment rulemaking, the primary mechanism through which citizens influence federal regulation. Whether these systems treat all public input equally remains largely untested. Using a counterfactual design, we held comment content constant and varied only the commenter's demographic attribution -- race, gender, and socioeconomic status -- to test whether eight LLMs available for federal use produce differential summaries of identical comments. We processed 182 public comments across 32 identity conditions, generating over 106,000 summaries. Occupation was the only identity signal to produce consistent differential treatment: the same comment attributed to a street vendor, compared to a financial analyst, received a summary that preserved less of the original meaning, used simpler language, and shifted emotional tone. This pattern held across all names, prompts, models, and regulatory contexts tested. Race effects were inconsistent and appeared driven by specific name tokens rather than racial categories; gender effects were absent. Writing quality predicted summarization outcomes through argument substance rather than surface mechanics; experimentally injected spelling and grammar errors had negligible effects. The magnitude of occupation-based differential treatment varied by model provider, meaning that selecting a model implicitly selects a level of fairness -- a dimension that current procurement frameworks such as FedRAMP do not evaluate. These findings suggest that socioeconomic signals warrant attention in AI fairness assessments for government information systems, and that fairness benchmarks could be incorporated into existing federal IT procurement processes.
title All Public Voices Are Equal, But Are Some More Equal Than Others to LLMs?
topic Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2604.17247