Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Lerner, Emilia Agis, Dorner, Florian E., Ash, Elliott, Goel, Naman |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Post-Processing-Based Fair Federated Learning Framework
by: Zhou, Yi, et al.
Published: (2025)
by: Zhou, Yi, et al.
Published: (2025)
Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
by: Barman, Kristian González, et al.
Published: (2024)
by: Barman, Kristian González, et al.
Published: (2024)
FairTargetSim: An Interactive Simulator for Understanding and Explaining the Fairness Effects of Target Variable Definition
by: Gala, Dalia, et al.
Published: (2024)
by: Gala, Dalia, et al.
Published: (2024)
Evaluating LLM Behavior in Hiring: Implicit Weights, Fairness Across Groups, and Alignment with Human Preferences
by: Hoffmann, Morgane, et al.
Published: (2026)
by: Hoffmann, Morgane, et al.
Published: (2026)
Fairness of ChatGPT
by: Li, Yunqi, et al.
Published: (2023)
by: Li, Yunqi, et al.
Published: (2023)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
by: Zhou, Han, et al.
Published: (2024)
by: Zhou, Han, et al.
Published: (2024)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers
by: Chakrabarty, Tuhin, et al.
Published: (2025)
by: Chakrabarty, Tuhin, et al.
Published: (2025)
Bias and Fairness in Large Language Models: A Survey
by: Gallegos, Isabel O., et al.
Published: (2023)
by: Gallegos, Isabel O., et al.
Published: (2023)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
by: Shieh, Evan, et al.
Published: (2024)
by: Shieh, Evan, et al.
Published: (2024)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
by: Deng, Wenlong, et al.
Published: (2024)
by: Deng, Wenlong, et al.
Published: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
Exploring Accuracy-Fairness Trade-off in Large Language Models
by: Zhang, Qingquan, et al.
Published: (2024)
by: Zhang, Qingquan, et al.
Published: (2024)
Measuring Political Preferences in AI Systems: An Integrative Approach
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Effect of Gender Fair Job Description on Generative AI Images
by: Böckling, Finn, et al.
Published: (2025)
by: Böckling, Finn, et al.
Published: (2025)
A Unifying Human-Centered AI Fairness Framework
by: Rahman, Munshi Mahbubur, et al.
Published: (2025)
by: Rahman, Munshi Mahbubur, et al.
Published: (2025)
AXOLOTL: Fairness through Assisted Self-Debiasing of Large Language Model Outputs
by: Ebrahimi, Sana, et al.
Published: (2024)
by: Ebrahimi, Sana, et al.
Published: (2024)
The Political Preferences of LLMs
by: Rozado, David
Published: (2024)
by: Rozado, David
Published: (2024)
Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models
by: Song, Kefan, et al.
Published: (2025)
by: Song, Kefan, et al.
Published: (2025)
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
by: Goel, Naman
Published: (2023)
by: Goel, Naman
Published: (2023)
First-Person Fairness in Chatbots
by: Eloundou, Tyna, et al.
Published: (2024)
by: Eloundou, Tyna, et al.
Published: (2024)
Stairway to Fairness: Connecting Group and Individual Fairness
by: Rampisela, Theresia Veronika, et al.
Published: (2025)
by: Rampisela, Theresia Veronika, et al.
Published: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Cost Efficient Fairness Audit Under Partial Feedback
by: Das, Nirjhar, et al.
Published: (2025)
by: Das, Nirjhar, et al.
Published: (2025)
Human Preferences for Constructive Interactions in Language Model Alignment
by: Kyrychenko, Yara, et al.
Published: (2025)
by: Kyrychenko, Yara, et al.
Published: (2025)
Misaligned by Reward: Socially Undesirable Preferences in LLMs
by: Ghazaryan, Gayane, et al.
Published: (2026)
by: Ghazaryan, Gayane, et al.
Published: (2026)
FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods
by: Han, Xiaotian, et al.
Published: (2023)
by: Han, Xiaotian, et al.
Published: (2023)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
Fairness in Federated Learning: Fairness for Whom?
by: Taik, Afaf, et al.
Published: (2025)
by: Taik, Afaf, et al.
Published: (2025)
On The Fairness Impacts of Hardware Selection in Machine Learning
by: Nelaturu, Sree Harsha, et al.
Published: (2023)
by: Nelaturu, Sree Harsha, et al.
Published: (2023)
Fair Classification with Partial Feedback: An Exploration-Based Data Collection Approach
by: Keswani, Vijay, et al.
Published: (2024)
by: Keswani, Vijay, et al.
Published: (2024)
Temporal Preferences in Language Models for Long-Horizon Assistance
by: Mazyaki, Ali, et al.
Published: (2025)
by: Mazyaki, Ali, et al.
Published: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
Mapping the Potential of Explainable AI for Fairness Along the AI Lifecycle
by: Deck, Luca, et al.
Published: (2024)
by: Deck, Luca, et al.
Published: (2024)
Impact of Fairness Regulations on Institutions' Policies and Population Qualifications
by: Montaseri, Hamidreza, et al.
Published: (2024)
by: Montaseri, Hamidreza, et al.
Published: (2024)
OxonFair: A Flexible Toolkit for Algorithmic Fairness
by: Delaney, Eoin, et al.
Published: (2024)
by: Delaney, Eoin, et al.
Published: (2024)
FairMT: Fairness for Heterogeneous Multi-Task Learning
by: Hu, Guanyu, et al.
Published: (2025)
by: Hu, Guanyu, et al.
Published: (2025)
Similar Items
-
A Post-Processing-Based Fair Federated Learning Framework
by: Zhou, Yi, et al.
Published: (2025) -
Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
by: Barman, Kristian González, et al.
Published: (2024) -
FairTargetSim: An Interactive Simulator for Understanding and Explaining the Fairness Effects of Target Variable Definition
by: Gala, Dalia, et al.
Published: (2024) -
Evaluating LLM Behavior in Hiring: Implicit Weights, Fairness Across Groups, and Alignment with Human Preferences
by: Hoffmann, Morgane, et al.
Published: (2026) -
Fairness of ChatGPT
by: Li, Yunqi, et al.
Published: (2023)