Evaluating Large Language Models for Fair and Reliable Organ Allocation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Brian Hyeongseok, Murray, Hannah, Lee, Isabelle, Byun, Jason, Lum, Joshua, Yogatama, Dani, Micha, Evi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
by: Lee, Isabelle, et al.
Published: (2024)
by: Lee, Isabelle, et al.
Published: (2024)
Group Fairness in Peer Review
by: Aziz, Haris, et al.
Published: (2024)
by: Aziz, Haris, et al.
Published: (2024)
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale
by: Lee, Isabelle, et al.
Published: (2025)
by: Lee, Isabelle, et al.
Published: (2025)
Rigorous Interpretation Is a Form of Evaluation
by: Lee, Isabelle, et al.
Published: (2026)
by: Lee, Isabelle, et al.
Published: (2026)
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
by: Chiang, Ting-Rui, et al.
Published: (2024)
by: Chiang, Ting-Rui, et al.
Published: (2024)
LocateBench: Evaluating the Locating Ability of Vision Language Models
by: Chiang, Ting-Rui, et al.
Published: (2024)
by: Chiang, Ting-Rui, et al.
Published: (2024)
Large Language Models for Interpretable Mental Health Diagnosis
by: Kim, Brian Hyeongseok, et al.
Published: (2025)
by: Kim, Brian Hyeongseok, et al.
Published: (2025)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024)
by: Liu, Ollie, et al.
Published: (2024)
Proportional Fairness in Non-Centroid Clustering
by: Caragiannis, Ioannis, et al.
Published: (2024)
by: Caragiannis, Ioannis, et al.
Published: (2024)
Evaluating Language Models for Harmful Manipulation
by: Akbulut, Canfer, et al.
Published: (2026)
by: Akbulut, Canfer, et al.
Published: (2026)
Fairness Testing of Large Language Models in Role-Playing
by: Li, Xinyue, et al.
Published: (2024)
by: Li, Xinyue, et al.
Published: (2024)
Can Large Language Models Develop Gambling Addiction?
by: Lee, Seungpil, et al.
Published: (2025)
by: Lee, Seungpil, et al.
Published: (2025)
Computing Voting Rules with Improvement Feedback
by: Micha, Evi, et al.
Published: (2025)
by: Micha, Evi, et al.
Published: (2025)
Bias and Fairness in Large Language Models: A Survey
by: Gallegos, Isabel O., et al.
Published: (2023)
by: Gallegos, Isabel O., et al.
Published: (2023)
Semantic Consistency for Assuring Reliability of Large Language Models
by: Raj, Harsh, et al.
Published: (2023)
by: Raj, Harsh, et al.
Published: (2023)
Deployment and Evaluation of an EHR-integrated, Large Language Model-Powered Tool to Triage Surgical Patients
by: Wang, Jane, et al.
Published: (2026)
by: Wang, Jane, et al.
Published: (2026)
Assessing the Reliability of Large Language Models in the Bengali Legal Context: A Comparative Evaluation Using LLM-as-Judge and Legal Experts
by: Aftahee, Sabik, et al.
Published: (2025)
by: Aftahee, Sabik, et al.
Published: (2025)
PerFairX: Is There a Balance Between Fairness and Personality in Large Language Model Recommendations?
by: Sah, Chandan Kumar
Published: (2025)
by: Sah, Chandan Kumar
Published: (2025)
Existential Conversations with Large Language Models: Content, Community, and Culture
by: Shanahan, Murray, et al.
Published: (2024)
by: Shanahan, Murray, et al.
Published: (2024)
Assessing the Reliability and Validity of Large Language Models for Automated Assessment of Student Essays in Higher Education
by: Gaggioli, Andrea, et al.
Published: (2025)
by: Gaggioli, Andrea, et al.
Published: (2025)
Temporal Panel Selection in Ongoing Citizens' Assemblies
by: Kalayci, Yusuf Hakan, et al.
Published: (2026)
by: Kalayci, Yusuf Hakan, et al.
Published: (2026)
Standing on FURM ground -- A framework for evaluating Fair, Useful, and Reliable AI Models in healthcare systems
by: Callahan, Alison, et al.
Published: (2024)
by: Callahan, Alison, et al.
Published: (2024)
Manipulation and the AI Act: Large Language Model Chatbots and the Danger of Mirrors
by: Krook, Joshua
Published: (2025)
by: Krook, Joshua
Published: (2025)
INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models
by: Jin, Di, et al.
Published: (2024)
by: Jin, Di, et al.
Published: (2024)
Large Language Models as Partners in Student Essay Evaluation
by: Ishida, Toru, et al.
Published: (2024)
by: Ishida, Toru, et al.
Published: (2024)
Exploring Accuracy-Fairness Trade-off in Large Language Models
by: Zhang, Qingquan, et al.
Published: (2024)
by: Zhang, Qingquan, et al.
Published: (2024)
On Retrieval Augmentation and the Limitations of Language Model Training
by: Chiang, Ting-Rui, et al.
Published: (2023)
by: Chiang, Ting-Rui, et al.
Published: (2023)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
by: Wang, Ze, et al.
Published: (2024)
by: Wang, Ze, et al.
Published: (2024)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
Normative Evaluation of Large Language Models with Everyday Moral Dilemmas
by: Sachdeva, Pratik S., et al.
Published: (2025)
by: Sachdeva, Pratik S., et al.
Published: (2025)
Evaluation of Bias Towards Medical Professionals in Large Language Models
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
On Fairness of Low-Rank Adaptation of Large Models
by: Ding, Zhoujie, et al.
Published: (2024)
by: Ding, Zhoujie, et al.
Published: (2024)
The Intersectionality Problem for Algorithmic Fairness
by: Himmelreich, Johannes, et al.
Published: (2024)
by: Himmelreich, Johannes, et al.
Published: (2024)
REQUAL-LM: Reliability and Equity through Aggregation in Large Language Models
by: Ebrahimi, Sana, et al.
Published: (2024)
by: Ebrahimi, Sana, et al.
Published: (2024)
On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts
by: Aremu, Toluwani, et al.
Published: (2024)
by: Aremu, Toluwani, et al.
Published: (2024)
Evaluating Large Language Models for Detecting Antisemitism
by: Patel, Jay, et al.
Published: (2025)
by: Patel, Jay, et al.
Published: (2025)
Evaluating Psychological Safety of Large Language Models
by: Li, Xingxuan, et al.
Published: (2022)
by: Li, Xingxuan, et al.
Published: (2022)
Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models
by: Song, Kefan, et al.
Published: (2025)
by: Song, Kefan, et al.
Published: (2025)
Fairness Evaluation for Uplift Modeling in the Absence of Ground Truth
by: Kadioglu, Serdar, et al.
Published: (2024)
by: Kadioglu, Serdar, et al.
Published: (2024)
Large Language Models as Misleading Assistants in Conversation
by: Hou, Betty Li, et al.
Published: (2024)
by: Hou, Betty Li, et al.
Published: (2024)
Similar Items
-
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
by: Lee, Isabelle, et al.
Published: (2024) -
Group Fairness in Peer Review
by: Aziz, Haris, et al.
Published: (2024) -
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale
by: Lee, Isabelle, et al.
Published: (2025) -
Rigorous Interpretation Is a Form of Evaluation
by: Lee, Isabelle, et al.
Published: (2026) -
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
by: Chiang, Ting-Rui, et al.
Published: (2024)