Human-Centered Design Recommendations for LLM-as-a-Judge
Fuente:
arXiv
Guardado en:
| Autores principales: | Pan, Qian, Ashktorab, Zahra, Desmond, Michael, Cooper, Martin Santillan, Johnson, James, Nair, Rahul, Daly, Elizabeth, Geyer, Werner |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
por: Ashktorab, Zahra, et al.
Publicado: (2025)
por: Ashktorab, Zahra, et al.
Publicado: (2025)
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
por: Ashktorab, Zahra, et al.
Publicado: (2024)
por: Ashktorab, Zahra, et al.
Publicado: (2024)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
por: Do, Hyo Jin, et al.
Publicado: (2025)
por: Do, Hyo Jin, et al.
Publicado: (2025)
Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations
por: Ashktorab, Zahra, et al.
Publicado: (2024)
por: Ashktorab, Zahra, et al.
Publicado: (2024)
Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
por: Song, Jaeyoon, et al.
Publicado: (2025)
por: Song, Jaeyoon, et al.
Publicado: (2025)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
por: Wagner, Nico, et al.
Publicado: (2024)
por: Wagner, Nico, et al.
Publicado: (2024)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
por: Chiang, Charles, et al.
Publicado: (2026)
por: Chiang, Charles, et al.
Publicado: (2026)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
por: Gebreegziabher, Simret Araya, et al.
Publicado: (2026)
por: Gebreegziabher, Simret Araya, et al.
Publicado: (2026)
Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability
por: Song, Jaeyoon, et al.
Publicado: (2025)
por: Song, Jaeyoon, et al.
Publicado: (2025)
Helping the Helper: Supporting Peer Counselors via AI-Empowered Practice and Feedback
por: Hsu, Shang-Ling, et al.
Publicado: (2023)
por: Hsu, Shang-Ling, et al.
Publicado: (2023)
A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development
por: Geyer, Werner, et al.
Publicado: (2025)
por: Geyer, Werner, et al.
Publicado: (2025)
Humble AI in the real-world: the case of algorithmic hiring
por: Nair, Rahul, et al.
Publicado: (2025)
por: Nair, Rahul, et al.
Publicado: (2025)
Who's Sorry Now: User Preferences Among Rote, Empathic, and Explanatory Apologies from LLM Chatbots
por: Ashktorab, Zahra, et al.
Publicado: (2025)
por: Ashktorab, Zahra, et al.
Publicado: (2025)
DesignerlyLoop: Forming Design Intent through Curated Reasoning for Human-LLM Alignment
por: Wang, Anqi, et al.
Publicado: (2025)
por: Wang, Anqi, et al.
Publicado: (2025)
Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
por: Do, Hyo Jin, et al.
Publicado: (2025)
por: Do, Hyo Jin, et al.
Publicado: (2025)
Design Principles for Generative AI Applications
por: Weisz, Justin D., et al.
Publicado: (2024)
por: Weisz, Justin D., et al.
Publicado: (2024)
Facilitating Human-LLM Collaboration through Factuality Scores and Source Attributions
por: Do, Hyo Jin, et al.
Publicado: (2024)
por: Do, Hyo Jin, et al.
Publicado: (2024)
Toward Humanity-Centered Design without Hubris
por: Gorichanaz, Tim
Publicado: (2024)
por: Gorichanaz, Tim
Publicado: (2024)
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
por: Nogueira, Brenda, et al.
Publicado: (2025)
por: Nogueira, Brenda, et al.
Publicado: (2025)
Human-Centered Design and Evaluation of a Workplace for the Remote Assistance of Highly Automated Vehicles
por: Schrank, Andreas, et al.
Publicado: (2023)
por: Schrank, Andreas, et al.
Publicado: (2023)
Current and Future Use of Large Language Models for Knowledge Work
por: Brachman, Michelle, et al.
Publicado: (2025)
por: Brachman, Michelle, et al.
Publicado: (2025)
Toward Human-Centered Human-AI Interaction: Advances in Theoretical Frameworks and Practice
por: Gao, Zaifeng, et al.
Publicado: (2026)
por: Gao, Zaifeng, et al.
Publicado: (2026)
Human-Centered Artificial Social Intelligence (HC-ASI)
por: Pan, Hanxi, et al.
Publicado: (2025)
por: Pan, Hanxi, et al.
Publicado: (2025)
ColorCode: A Bayesian Approach to Augmentative and Alternative Communication with Two Buttons
por: Daly, Matthew
Publicado: (2022)
por: Daly, Matthew
Publicado: (2022)
Organizational Practices and Socio-Technical Design of Human-Centered AI
por: Herrmann, Thomas
Publicado: (2026)
por: Herrmann, Thomas
Publicado: (2026)
Designing Around Stigma: Human-Centered LLMs for Menstrual Health
por: Shahnawaz, Amna, et al.
Publicado: (2026)
por: Shahnawaz, Amna, et al.
Publicado: (2026)
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
por: Do, Hyo Jin, et al.
Publicado: (2025)
por: Do, Hyo Jin, et al.
Publicado: (2025)
Applying LLM-Powered Virtual Humans to Child Interviews in Child-Centered Design
por: Li, Linshi, et al.
Publicado: (2025)
por: Li, Linshi, et al.
Publicado: (2025)
Identifying the Barriers to Human-Centered Design in the Workplace: Perspectives from UX Professionals
por: Gorichanaz, Tim
Publicado: (2024)
por: Gorichanaz, Tim
Publicado: (2024)
Human-Centered AI Maturity Model (HCAI-MM): An Organizational Design Perspective
por: Winby, Stuart, et al.
Publicado: (2025)
por: Winby, Stuart, et al.
Publicado: (2025)
Toward Needs-Conscious Design: Co-Designing a Human-Centered Framework for AI-Mediated Communication
por: Wolfe, Robert, et al.
Publicado: (2025)
por: Wolfe, Robert, et al.
Publicado: (2025)
Beyond Code Generation: LLM-supported Exploration of the Program Design Space
por: Zamfirescu-Pereira, J. D., et al.
Publicado: (2025)
por: Zamfirescu-Pereira, J. D., et al.
Publicado: (2025)
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
por: Szymanski, Annalisa, et al.
Publicado: (2024)
por: Szymanski, Annalisa, et al.
Publicado: (2024)
CogInstrument: Modeling Cognitive Processes for Bidirectional Human-LLM Alignment in Planning Tasks
por: Wang, Anqi, et al.
Publicado: (2026)
por: Wang, Anqi, et al.
Publicado: (2026)
Environment-Aware and Human-Cooperative Swing Control for Lower-Limb Prostheses in Diverse Obstacle Scenarios
por: Xing, Haosen, et al.
Publicado: (2025)
por: Xing, Haosen, et al.
Publicado: (2025)
"The Diagram is like Guardrails": Structuring GenAI-assisted Hypotheses Exploration with an Interactive Shared Representation
por: Ding, Zijian, et al.
Publicado: (2025)
por: Ding, Zijian, et al.
Publicado: (2025)
Designing an intelligent computer game for predicting dysgraphia
por: Nevisi, Zahra, et al.
Publicado: (2025)
por: Nevisi, Zahra, et al.
Publicado: (2025)
Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
por: Qian, Crystal, et al.
Publicado: (2024)
por: Qian, Crystal, et al.
Publicado: (2024)
Large Language Model Agent Personality and Response Appropriateness: Evaluation by Human Linguistic Experts, LLM-as-Judge, and Natural Language Processing Model
por: Jayakumar, Eswari, et al.
Publicado: (2025)
por: Jayakumar, Eswari, et al.
Publicado: (2025)
Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care
por: Shukla, Shivam, et al.
Publicado: (2026)
por: Shukla, Shivam, et al.
Publicado: (2026)
Ejemplares similares
-
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
por: Ashktorab, Zahra, et al.
Publicado: (2025) -
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
por: Ashktorab, Zahra, et al.
Publicado: (2024) -
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
por: Do, Hyo Jin, et al.
Publicado: (2025) -
Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations
por: Ashktorab, Zahra, et al.
Publicado: (2024) -
Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
por: Song, Jaeyoon, et al.
Publicado: (2025)