Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yinong Oliver, Sivakumar, Nivedha, Khan, Falaah Arif, Susa, Rin Metcalf, Golinski, Adam, Mackraz, Natalie, Theobald, Barry-John, Zappella, Luca, Apostoloff, Nicholas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
by: Khan, Falaah Arif, et al.
Published: (2025)
by: Khan, Falaah Arif, et al.
Published: (2025)
Fairness Dynamics During Training
by: Patel, Krishna, et al.
Published: (2025)
by: Patel, Krishna, et al.
Published: (2025)
Bias after Prompting: Persistent Discrimination in Large Language Models
by: Sivakumar, Nivedha, et al.
Published: (2025)
by: Sivakumar, Nivedha, et al.
Published: (2025)
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
by: Mackraz, Natalie, et al.
Published: (2024)
by: Mackraz, Natalie, et al.
Published: (2024)
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)
by: Paes, Lucas Monteiro, et al.
Published: (2025)
Aligning LLMs by Predicting Preferences from User Writing Samples
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards
by: Metcalf, Katherine, et al.
Published: (2024)
by: Metcalf, Katherine, et al.
Published: (2024)
PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories
by: Aroca-Ouellette, Stephane, et al.
Published: (2024)
by: Aroca-Ouellette, Stephane, et al.
Published: (2024)
Theoretical Limits of Language Model Alignment
by: Paes, Lucas Monteiro, et al.
Published: (2026)
by: Paes, Lucas Monteiro, et al.
Published: (2026)
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
by: Suau, Xavier, et al.
Published: (2024)
by: Suau, Xavier, et al.
Published: (2024)
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025)
by: Santilli, Andrea, et al.
Published: (2025)
Controlling Language and Diffusion Models by Transporting Activations
by: Rodriguez, Pau, et al.
Published: (2024)
by: Rodriguez, Pau, et al.
Published: (2024)
An Epistemic and Aleatoric Decomposition of Arbitrariness to Constrain the Set of Good Models
by: Khan, Falaah Arif, et al.
Published: (2023)
by: Khan, Falaah Arif, et al.
Published: (2023)
Still More Shades of Null: An Evaluation Suite for Responsible Missing Value Imputation
by: Khan, Falaah Arif, et al.
Published: (2024)
by: Khan, Falaah Arif, et al.
Published: (2024)
Considerations for Distribution Shift Robustness of Diagnostic Models in Healthcare
by: Blaas, Arno, et al.
Published: (2024)
by: Blaas, Arno, et al.
Published: (2024)
With a Grain of SALT: Are LLMs Fair Across Social Dimensions?
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
STRUCTURAL FRAGMENTATION IN INDIAN ENVIRONMENTAL LAW: ENFORCEMENT DEFICITS AND THE IMPERATIVE FOR A UNIFIED ENVIRONMENTAL CODE
by: Joan Nivedha S
Published: (2026)
by: Joan Nivedha S
Published: (2026)
LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs
by: Chen, Chacha, et al.
Published: (2026)
by: Chen, Chacha, et al.
Published: (2026)
FairlyUncertain: A Comprehensive Benchmark of Uncertainty in Algorithmic Fairness
by: Rosenblatt, Lucas, et al.
Published: (2024)
by: Rosenblatt, Lucas, et al.
Published: (2024)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
by: Mao, Hanyi, et al.
Published: (2025)
by: Mao, Hanyi, et al.
Published: (2025)
DS FedProxGrad: Asymptotic Stationarity Without Noise Floor in Fair Federated Learning
by: Arif, Huzaifa
Published: (2025)
by: Arif, Huzaifa
Published: (2025)
Counterfactual Fairness with Graph Uncertainty
by: Valério, Davi, et al.
Published: (2026)
by: Valério, Davi, et al.
Published: (2026)
Uncertainty-based Fairness Measures
by: Kuzucu, Selim, et al.
Published: (2023)
by: Kuzucu, Selim, et al.
Published: (2023)
The Impossibility of Fair LLMs
by: Anthis, Jacy, et al.
Published: (2024)
by: Anthis, Jacy, et al.
Published: (2024)
Wearable-based Fair and Accurate Pain Assessment Using Multi-Attribute Fairness Loss in Convolutional Neural Networks
by: Zhu, Yidong, et al.
Published: (2023)
by: Zhu, Yidong, et al.
Published: (2023)
Fair-Gate: Fairness-Aware Interpretable Risk Gating for Sex-Fair Voice Biometrics
by: Qu, Yangyang, et al.
Published: (2026)
by: Qu, Yangyang, et al.
Published: (2026)
We Are AI: Taking Control of Technology
by: Stoyanovich, Julia, et al.
Published: (2025)
by: Stoyanovich, Julia, et al.
Published: (2025)
(Un)certainty of (Un)fairness: Preference-Based Selection of Certainly Fair Decision-Makers
by: Duong, Manh Khoi, et al.
Published: (2024)
by: Duong, Manh Khoi, et al.
Published: (2024)
Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs
by: Bouchard, Dylan
Published: (2024)
by: Bouchard, Dylan
Published: (2024)
A note on precotangent bundles: the example of Grassmannians
by: Goliński, Tomasz
Published: (2025)
by: Goliński, Tomasz
Published: (2025)
How Fair is Your Diffusion Recommender Model?
by: Malitesta, Daniele, et al.
Published: (2024)
by: Malitesta, Daniele, et al.
Published: (2024)
Fair Allocation with Money: What is Your Objective?
by: Elmalem, Noga Klein, et al.
Published: (2025)
by: Elmalem, Noga Klein, et al.
Published: (2025)
Achievable Fairness on Your Data With Utility Guarantees
by: Taufiq, Muhammad Faaiz, et al.
Published: (2024)
by: Taufiq, Muhammad Faaiz, et al.
Published: (2024)
ExpertLens: Activation steering features are highly interpretable
by: Fedzechkina, Masha, et al.
Published: (2025)
by: Fedzechkina, Masha, et al.
Published: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
by: Sundar, Anirudh, et al.
Published: (2025)
by: Sundar, Anirudh, et al.
Published: (2025)
The Illusion of Fairness: Auditing Fairness Interventions with Audit Studies
by: Sariola, Disa, et al.
Published: (2025)
by: Sariola, Disa, et al.
Published: (2025)
Are Your Models Still Fair? Fairness Attacks on Graph Neural Networks via Node Injections
by: Luo, Zihan, et al.
Published: (2024)
by: Luo, Zihan, et al.
Published: (2024)
Fairness in Ranking under Disparate Uncertainty
by: Rastogi, Richa, et al.
Published: (2023)
by: Rastogi, Richa, et al.
Published: (2023)
Fair Uncertainty Quantification for Depression Prediction
by: Li, Yonghong, et al.
Published: (2025)
by: Li, Yonghong, et al.
Published: (2025)
Similar Items
-
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
by: Khan, Falaah Arif, et al.
Published: (2025) -
Fairness Dynamics During Training
by: Patel, Krishna, et al.
Published: (2025) -
Bias after Prompting: Persistent Discrimination in Large Language Models
by: Sivakumar, Nivedha, et al.
Published: (2025) -
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
by: Mackraz, Natalie, et al.
Published: (2024) -
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)