Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
Fuente:
arXiv
Saved in:
| Main Authors: | Pfohl, Stephen R., Harris, Natalie, Nagpal, Chirag, Madras, David, Mhasawade, Vishwali, Salaudeen, Olawale, Dieng, Awa, Sequeira, Shannon, Arciniegas, Santiago, Sung, Lillian, Ezeanochie, Nnamdi, Cole-Lewis, Heather, Heller, Katherine, Koyejo, Sanmi, D'Amour, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causally Inspired Regularization Enables Domain General Representations
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
Proxy Methods for Domain Adaptation
by: Tsai, Katherine, et al.
Published: (2024)
by: Tsai, Katherine, et al.
Published: (2024)
The Case for Globalizing Fairness: A Mixed Methods Study on Colonialism, AI, and Health in Africa
by: Asiedu, Mercy, et al.
Published: (2024)
by: Asiedu, Mercy, et al.
Published: (2024)
Transforming and Combining Rewards for Aligning Large Language Models
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Disparate Effect Of Missing Mediators On Transportability of Causal Effects
by: Mhasawade, Vishwali, et al.
Published: (2024)
by: Mhasawade, Vishwali, et al.
Published: (2024)
Globalizing Fairness Attributes in Machine Learning: A Case Study on Health in Africa
by: Asiedu, Mercy Nyamewaa, et al.
Published: (2023)
by: Asiedu, Mercy Nyamewaa, et al.
Published: (2023)
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
by: Lum, Kristian, et al.
Published: (2024)
by: Lum, Kristian, et al.
Published: (2024)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023)
by: Eisenstein, Jacob, et al.
Published: (2023)
Impact on Public Health Decision Making by Utilizing Big Data Without Domain Knowledge
by: Zhang, Miao, et al.
Published: (2024)
by: Zhang, Miao, et al.
Published: (2024)
ImageNot: A contrast with ImageNet preserves model rankings
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
Understanding Disparities in Post Hoc Machine Learning Explanation
by: Mhasawade, Vishwali, et al.
Published: (2024)
by: Mhasawade, Vishwali, et al.
Published: (2024)
Theoretical guarantees on the best-of-n alignment policy
by: Beirami, Ahmad, et al.
Published: (2024)
by: Beirami, Ahmad, et al.
Published: (2024)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
by: Zheng, Jiajing, et al.
Published: (2021)
by: Zheng, Jiajing, et al.
Published: (2021)
Preference Models assume Proportional Hazards of Utilities
by: Nagpal, Chirag
Published: (2025)
by: Nagpal, Chirag
Published: (2025)
Nteasee: Understanding Needs in AI for Health in Africa -- A Mixed-Methods Study of Expert and General Population Perspectives
by: Asiedu, Mercy Nyamewaa, et al.
Published: (2024)
by: Asiedu, Mercy Nyamewaa, et al.
Published: (2024)
Toward an Evaluation Science for Generative AI Systems
by: Weidinger, Laura, et al.
Published: (2025)
by: Weidinger, Laura, et al.
Published: (2025)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
by: Robertson, Zachary, et al.
Published: (2025)
by: Robertson, Zachary, et al.
Published: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
A Framework for Objective-Driven Dynamical Stochastic Fields
by: Zhang, Yibo Jacky, et al.
Published: (2025)
by: Zhang, Yibo Jacky, et al.
Published: (2025)
Choosing a Proxy Metric from Past Experiments
by: Tripuraneni, Nilesh, et al.
Published: (2023)
by: Tripuraneni, Nilesh, et al.
Published: (2023)
Making physical activity fun and accessible to adults with intellectual disabilities: A pilot study of a gamification intervention
by: Stéphanie Turgeon, et al.
Published: (2024)
by: Stéphanie Turgeon, et al.
Published: (2024)
Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
What's in a Query: Polarity-Aware Distribution-Based Fair Ranking
by: Balagopalan, Aparna, et al.
Published: (2025)
by: Balagopalan, Aparna, et al.
Published: (2025)
In-Context Learning of Energy Functions
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Discovering Implicit Large Language Model Alignment Objectives
by: Chen, Edward, et al.
Published: (2026)
by: Chen, Edward, et al.
Published: (2026)
High-Dimensional Markov-switching Ordinary Differential Processes
by: Tsai, Katherine, et al.
Published: (2024)
by: Tsai, Katherine, et al.
Published: (2024)
Distributional Machine Unlearning via Selective Data Removal
by: Allouah, Youssef, et al.
Published: (2025)
by: Allouah, Youssef, et al.
Published: (2025)
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
by: Zhu, Junzhe, et al.
Published: (2023)
by: Zhu, Junzhe, et al.
Published: (2023)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
by: Nagpal, Chirag, et al.
Published: (2024)
by: Nagpal, Chirag, et al.
Published: (2024)
Machine Learning for Health symposium 2024 -- Findings track
by: Hegselmann, Stefan, et al.
Published: (2025)
by: Hegselmann, Stefan, et al.
Published: (2025)
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
by: Sühr, Tom, et al.
Published: (2025)
by: Sühr, Tom, et al.
Published: (2025)
Motetten am Hof Maximilians II. (1527–1576)
by: Pfohl, Jonas
Published: (2022)
by: Pfohl, Jonas
Published: (2022)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency
by: Zhang, Yibo Jacky, et al.
Published: (2026)
by: Zhang, Yibo Jacky, et al.
Published: (2026)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
by: Wang, Angelina, et al.
Published: (2025)
by: Wang, Angelina, et al.
Published: (2025)
Principled Federated Domain Adaptation: Gradient Projection and Auto-Weighting
by: Jiang, Enyi, et al.
Published: (2023)
by: Jiang, Enyi, et al.
Published: (2023)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Similar Items
-
Causally Inspired Regularization Enables Domain General Representations
by: Salaudeen, Olawale, et al.
Published: (2024) -
Proxy Methods for Domain Adaptation
by: Tsai, Katherine, et al.
Published: (2024) -
The Case for Globalizing Fairness: A Mixed Methods Study on Colonialism, AI, and Health in Africa
by: Asiedu, Mercy, et al.
Published: (2024) -
Transforming and Combining Rewards for Aligning Large Language Models
by: Wang, Zihao, et al.
Published: (2024) -
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
by: Salaudeen, Olawale, et al.
Published: (2025)