Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Petrova, Nora, Gordon, Andrew, Blindow, Enzo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating LLMs as Human Surrogates in Controlled Experiments
von: Hoq, Adnan, et al.
Veröffentlicht: (2026)
von: Hoq, Adnan, et al.
Veröffentlicht: (2026)
Aligning LLMs with Individual Preferences via Interaction
von: Wu, Shujin, et al.
Veröffentlicht: (2024)
von: Wu, Shujin, et al.
Veröffentlicht: (2024)
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
von: Daynauth, Roland, et al.
Veröffentlicht: (2024)
von: Daynauth, Roland, et al.
Veröffentlicht: (2024)
ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
von: Banyas, Peter, et al.
Veröffentlicht: (2025)
von: Banyas, Peter, et al.
Veröffentlicht: (2025)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
von: McCullum, Lucas, et al.
Veröffentlicht: (2025)
von: McCullum, Lucas, et al.
Veröffentlicht: (2025)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
von: Giorgi, Tommaso, et al.
Veröffentlicht: (2024)
von: Giorgi, Tommaso, et al.
Veröffentlicht: (2024)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
von: Shen, Hua, et al.
Veröffentlicht: (2024)
von: Shen, Hua, et al.
Veröffentlicht: (2024)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
von: Pang, Rock Yuren, et al.
Veröffentlicht: (2025)
von: Pang, Rock Yuren, et al.
Veröffentlicht: (2025)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
von: Sung, Yoo Yeon, et al.
Veröffentlicht: (2025)
von: Sung, Yoo Yeon, et al.
Veröffentlicht: (2025)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
von: Sun, Chongyan, et al.
Veröffentlicht: (2024)
von: Sun, Chongyan, et al.
Veröffentlicht: (2024)
HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?
von: Ji, Sijie, et al.
Veröffentlicht: (2024)
von: Ji, Sijie, et al.
Veröffentlicht: (2024)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
von: Soni, Nikita, et al.
Veröffentlicht: (2025)
von: Soni, Nikita, et al.
Veröffentlicht: (2025)
Towards Human-Centered RegTech: Unpacking Professionals' Strategies and Needs for Using LLMs Safely
von: Hu, Siying, et al.
Veröffentlicht: (2025)
von: Hu, Siying, et al.
Veröffentlicht: (2025)
Human Preferences for Constructive Interactions in Language Model Alignment
von: Kyrychenko, Yara, et al.
Veröffentlicht: (2025)
von: Kyrychenko, Yara, et al.
Veröffentlicht: (2025)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025)
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
von: Garbacea, Cristina, et al.
Veröffentlicht: (2026)
von: Garbacea, Cristina, et al.
Veröffentlicht: (2026)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
von: Isaza-Giraldo, Andrés, et al.
Veröffentlicht: (2025)
von: Isaza-Giraldo, Andrés, et al.
Veröffentlicht: (2025)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
von: Yadav, Neemesh, et al.
Veröffentlicht: (2025)
von: Yadav, Neemesh, et al.
Veröffentlicht: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
von: Chaleshtori, Fateme Hashemi, et al.
Veröffentlicht: (2024)
von: Chaleshtori, Fateme Hashemi, et al.
Veröffentlicht: (2024)
Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
von: Shen, Hua, et al.
Veröffentlicht: (2025)
von: Shen, Hua, et al.
Veröffentlicht: (2025)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
SAC: A Framework for Measuring and Inducing Personality Traits in LLMs with Dynamic Intensity Control
von: Chittem, Adithya, et al.
Veröffentlicht: (2025)
von: Chittem, Adithya, et al.
Veröffentlicht: (2025)
A Scalable Framework for Evaluating Health Language Models
von: Mallinar, Neil, et al.
Veröffentlicht: (2025)
von: Mallinar, Neil, et al.
Veröffentlicht: (2025)
Unpacking Interpretability: Human-Centered Criteria for Optimal Combinatorial Solutions
von: Pegler, Dominik, et al.
Veröffentlicht: (2026)
von: Pegler, Dominik, et al.
Veröffentlicht: (2026)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
von: Zeng, Qiuhai, et al.
Veröffentlicht: (2025)
von: Zeng, Qiuhai, et al.
Veröffentlicht: (2025)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
von: Huq, Faria, et al.
Veröffentlicht: (2025)
von: Huq, Faria, et al.
Veröffentlicht: (2025)
The Effectiveness of Style Vectors for Steering Large Language Models: A Human Evaluation
von: Diallo, Diaoulé, et al.
Veröffentlicht: (2026)
von: Diallo, Diaoulé, et al.
Veröffentlicht: (2026)
Using Generative Text Models to Create Qualitative Codebooks for Student Evaluations of Teaching
von: Katz, Andrew, et al.
Veröffentlicht: (2024)
von: Katz, Andrew, et al.
Veröffentlicht: (2024)
Human Evaluation of Procedural Knowledge Graph Extraction from Text with Large Language Models
von: Carriero, Valentina Anita, et al.
Veröffentlicht: (2024)
von: Carriero, Valentina Anita, et al.
Veröffentlicht: (2024)
Clinical knowledge in LLMs does not translate to human interactions
von: Bean, Andrew M., et al.
Veröffentlicht: (2025)
von: Bean, Andrew M., et al.
Veröffentlicht: (2025)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
von: Liu, Tianjian, et al.
Veröffentlicht: (2025)
von: Liu, Tianjian, et al.
Veröffentlicht: (2025)
Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases
von: Ford, Casey, et al.
Veröffentlicht: (2026)
von: Ford, Casey, et al.
Veröffentlicht: (2026)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
von: Lee, Suhyun, et al.
Veröffentlicht: (2026)
von: Lee, Suhyun, et al.
Veröffentlicht: (2026)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
von: Kwon, Deuksin, et al.
Veröffentlicht: (2025)
von: Kwon, Deuksin, et al.
Veröffentlicht: (2025)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
von: Sinacola, Enzo, et al.
Veröffentlicht: (2025)
von: Sinacola, Enzo, et al.
Veröffentlicht: (2025)
On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts
von: Aremu, Toluwani, et al.
Veröffentlicht: (2024)
von: Aremu, Toluwani, et al.
Veröffentlicht: (2024)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
von: Subedi, Krishna
Veröffentlicht: (2025)
von: Subedi, Krishna
Veröffentlicht: (2025)
Human-Centered AI in Multidisciplinary Medical Discussions: Evaluating the Feasibility of a Chat-Based Approach to Case Assessment
von: Sawano, Shinnosuke, et al.
Veröffentlicht: (2025)
von: Sawano, Shinnosuke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating LLMs as Human Surrogates in Controlled Experiments
von: Hoq, Adnan, et al.
Veröffentlicht: (2026) -
Aligning LLMs with Individual Preferences via Interaction
von: Wu, Shujin, et al.
Veröffentlicht: (2024) -
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
von: Daynauth, Roland, et al.
Veröffentlicht: (2024) -
ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
von: Banyas, Peter, et al.
Veröffentlicht: (2025) -
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
von: McCullum, Lucas, et al.
Veröffentlicht: (2025)