The Impossibility of Fair LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Anthis, Jacy, Lum, Kristian, Ekstrand, Michael, Feller, Avi, Tan, Chenhao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025)
by: Sturgeon, Benjamin, et al.
Published: (2025)
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
by: Lum, Kristian, et al.
Published: (2024)
by: Lum, Kristian, et al.
Published: (2024)
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
by: Garbacea, Cristina, et al.
Published: (2026)
by: Garbacea, Cristina, et al.
Published: (2026)
ALICE: Combining Feature Selection and Inter-Rater Agreeability for Machine Learning Insights
by: Anasashvili, Bachana, et al.
Published: (2024)
by: Anasashvili, Bachana, et al.
Published: (2024)
Data Quality in Crowdsourcing and Spamming Behavior Detection
by: Ba, Yang, et al.
Published: (2024)
by: Ba, Yang, et al.
Published: (2024)
Human-in-the-Loop Feature Selection Using Interpretable Kolmogorov-Arnold Network-based Double Deep Q-Network
by: Jahin, Md Abrar, et al.
Published: (2024)
by: Jahin, Md Abrar, et al.
Published: (2024)
Evaluating Imputation Techniques for Short-Term Gaps in Heart Rate Data
by: Gupta, Vaibhav, et al.
Published: (2025)
by: Gupta, Vaibhav, et al.
Published: (2025)
Tailored Behavior-Change Messaging for Physical Activity: Integrating Contextual Bandits and Large Language Models
by: Song, Haochen, et al.
Published: (2025)
by: Song, Haochen, et al.
Published: (2025)
Which Artificial Intelligences Do People Care About Most? A Conjoint Experiment on Moral Consideration
by: Ladak, Ali, et al.
Published: (2024)
by: Ladak, Ali, et al.
Published: (2024)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026)
by: Anthis, Jacy Reese, et al.
Published: (2026)
ClickTree: A Tree-based Method for Predicting Math Students' Performance Based on Clickstream Data
by: Rohani, Narjes, et al.
Published: (2024)
by: Rohani, Narjes, et al.
Published: (2024)
Robots, Chatbots, Self-Driving Cars: Perceptions of Mind and Morality Across Artificial Intelligences
by: Ladak, Ali, et al.
Published: (2025)
by: Ladak, Ali, et al.
Published: (2025)
The Dynamics of Delusion: Modeling Bidirectional False Belief Amplification in Human-Chatbot Dialogue
by: Mehta, Ashish, et al.
Published: (2026)
by: Mehta, Ashish, et al.
Published: (2026)
DataMap: A Portable Application for Visualizing High-Dimensional Data
by: Ge, Xijin
Published: (2025)
by: Ge, Xijin
Published: (2025)
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023)
by: Kıcıman, Emre, et al.
Published: (2023)
Forecasting Occupational Survivability of Rickshaw Pullers in a Changing Climate with Wearable Data
by: Rahaman, Masfiqur, et al.
Published: (2025)
by: Rahaman, Masfiqur, et al.
Published: (2025)
Crowdsourced Adaptive Surveys
by: Velez, Yamil
Published: (2024)
by: Velez, Yamil
Published: (2024)
Limits of Large Language Models in Debating Humans
by: Flamino, James, et al.
Published: (2024)
by: Flamino, James, et al.
Published: (2024)
Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
by: Hartmann, David, et al.
Published: (2024)
by: Hartmann, David, et al.
Published: (2024)
LLM Social Simulations Are a Promising Research Method
by: Anthis, Jacy Reese, et al.
Published: (2025)
by: Anthis, Jacy Reese, et al.
Published: (2025)
ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms
by: Vail, Alexandria K., et al.
Published: (2026)
by: Vail, Alexandria K., et al.
Published: (2026)
Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition
by: Feng, Kehua, et al.
Published: (2024)
by: Feng, Kehua, et al.
Published: (2024)
Harmonic LLMs are Trustworthy
by: Kersting, Nicholas S., et al.
Published: (2024)
by: Kersting, Nicholas S., et al.
Published: (2024)
Self-reflecting Large Language Models: A Hegelian Dialectical Approach
by: Abdali, Sara, et al.
Published: (2025)
by: Abdali, Sara, et al.
Published: (2025)
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
by: Hedderich, Michael A., et al.
Published: (2025)
by: Hedderich, Michael A., et al.
Published: (2025)
Processes Matter: How ML/GAI Approaches Could Support Open Qualitative Coding of Online Discourse Datasets
by: Chen, John, et al.
Published: (2025)
by: Chen, John, et al.
Published: (2025)
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
The AI Double Standard: Humans Judge All AIs for the Actions of One
by: Manoli, Aikaterina, et al.
Published: (2024)
by: Manoli, Aikaterina, et al.
Published: (2024)
Introducing MeMo: A Multimodal Dataset for Memory Modelling in Multiparty Conversations
by: Tsfasman, Maria, et al.
Published: (2024)
by: Tsfasman, Maria, et al.
Published: (2024)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
by: Jones, Graham M., et al.
Published: (2024)
by: Jones, Graham M., et al.
Published: (2024)
LLMs for XAI: Future Directions for Explaining Explanations
by: Zytek, Alexandra, et al.
Published: (2024)
by: Zytek, Alexandra, et al.
Published: (2024)
Evaluating the Usability of LLMs in Threat Intelligence Enrichment
by: Srikanth, Sanchana, et al.
Published: (2024)
by: Srikanth, Sanchana, et al.
Published: (2024)
Interaction Dynamics as a Reward Signal for LLMs
by: Gooding, Sian, et al.
Published: (2025)
by: Gooding, Sian, et al.
Published: (2025)
Generative UI: LLMs are Effective UI Generators
by: Leviathan, Yaniv, et al.
Published: (2026)
by: Leviathan, Yaniv, et al.
Published: (2026)
Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation
by: Ma, Cheng Charles, et al.
Published: (2024)
by: Ma, Cheng Charles, et al.
Published: (2024)
Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational Theory
by: Yuan, Yunhao, et al.
Published: (2025)
by: Yuan, Yunhao, et al.
Published: (2025)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
Model-in-the-Loop (MILO): Accelerating Multimodal AI Data Annotation with LLMs
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
by: Zeng, Qiuhai, et al.
Published: (2025)
by: Zeng, Qiuhai, et al.
Published: (2025)
Similar Items
-
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025) -
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
by: Lum, Kristian, et al.
Published: (2024) -
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
by: Garbacea, Cristina, et al.
Published: (2026) -
ALICE: Combining Feature Selection and Inter-Rater Agreeability for Machine Learning Insights
by: Anasashvili, Bachana, et al.
Published: (2024) -
Data Quality in Crowdsourcing and Spamming Behavior Detection
by: Ba, Yang, et al.
Published: (2024)