Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Szymanski, Annalisa, Ziems, Noah, Eicher-Miller, Heather A., Li, Toby Jia-Jun, Jiang, Meng, Metoyer, Ronald A. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
di: Szymanski, Annalisa, et al.
Pubblicazione: (2026)
di: Szymanski, Annalisa, et al.
Pubblicazione: (2026)
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024)
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
di: Chiang, Charles, et al.
Pubblicazione: (2026)
di: Chiang, Charles, et al.
Pubblicazione: (2026)
"I'm categorizing LLM as a productivity tool": Examining ethics of LLM use in HCI research practices
di: Kapania, Shivani, et al.
Pubblicazione: (2024)
di: Kapania, Shivani, et al.
Pubblicazione: (2024)
Large Language Model Agent Personality and Response Appropriateness: Evaluation by Human Linguistic Experts, LLM-as-Judge, and Natural Language Processing Model
di: Jayakumar, Eswari, et al.
Pubblicazione: (2025)
di: Jayakumar, Eswari, et al.
Pubblicazione: (2025)
EVOLVE: Emotion and Visual Output Learning via LLM Evaluation
di: Sinclair, Jordan, et al.
Pubblicazione: (2024)
di: Sinclair, Jordan, et al.
Pubblicazione: (2024)
My Favorite Streamer is an LLM: Discovering, Bonding, and Co-Creating in AI VTuber Fandom
di: Ye, Jiayi, et al.
Pubblicazione: (2025)
di: Ye, Jiayi, et al.
Pubblicazione: (2025)
Bridging the AI Adoption Gap: Designing an Interactive Pedagogical Agent for Higher Education Instructors
di: Chen, Si, et al.
Pubblicazione: (2025)
di: Chen, Si, et al.
Pubblicazione: (2025)
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
di: Shankar, Shreya, et al.
Pubblicazione: (2024)
di: Shankar, Shreya, et al.
Pubblicazione: (2024)
A Taxonomy for Human-LLM Interaction Modes: An Initial Exploration
di: Gao, Jie, et al.
Pubblicazione: (2024)
di: Gao, Jie, et al.
Pubblicazione: (2024)
Human-Centered Design Recommendations for LLM-as-a-Judge
di: Pan, Qian, et al.
Pubblicazione: (2024)
di: Pan, Qian, et al.
Pubblicazione: (2024)
Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code Modification
di: Tang, Ningzhi, et al.
Pubblicazione: (2025)
di: Tang, Ningzhi, et al.
Pubblicazione: (2025)
A Study on Developer Behaviors for Validating and Repairing LLM-Generated Code Using Eye Tracking and IDE Actions
di: Tang, Ningzhi, et al.
Pubblicazione: (2024)
di: Tang, Ningzhi, et al.
Pubblicazione: (2024)
CLEAR: Towards Contextual LLM-Empowered Privacy Policy Analysis and Risk Generation for Large Language Model Applications
di: Chen, Chaoran, et al.
Pubblicazione: (2024)
di: Chen, Chaoran, et al.
Pubblicazione: (2024)
Stayin' Aligned Over Time: Towards Longitudinal Human-LLM Alignment via Contextual Reflection and Privacy-Preserving Behavioral Data
di: Gebreegziabher, Simret Araya, et al.
Pubblicazione: (2026)
di: Gebreegziabher, Simret Araya, et al.
Pubblicazione: (2026)
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
di: Ashktorab, Zahra, et al.
Pubblicazione: (2025)
di: Ashktorab, Zahra, et al.
Pubblicazione: (2025)
An Empathy-Based Sandbox Approach to Bridge the Privacy Gap among Attitudes, Goals, Knowledge, and Behaviors
di: Chen, Chaoran, et al.
Pubblicazione: (2023)
di: Chen, Chaoran, et al.
Pubblicazione: (2023)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
di: Do, Hyo Jin, et al.
Pubblicazione: (2025)
di: Do, Hyo Jin, et al.
Pubblicazione: (2025)
ChartifyText: Automated Chart Generation from Data-Involved Texts via LLM
di: Zhang, Songheng, et al.
Pubblicazione: (2024)
di: Zhang, Songheng, et al.
Pubblicazione: (2024)
AI Academy: Building Generative AI Literacy in Higher Ed Instructors
di: Chen, Si, et al.
Pubblicazione: (2025)
di: Chen, Si, et al.
Pubblicazione: (2025)
GraphPilot: GUI Task Automation with One-Step LLM Reasoning Powered by Knowledge Graph
di: Yu, Mingxian, et al.
Pubblicazione: (2026)
di: Yu, Mingxian, et al.
Pubblicazione: (2026)
Knowledge Synthesis Graph: An LLM-Based Approach for Modeling Student Collaborative Discourse
di: Shui, Bo, et al.
Pubblicazione: (2026)
di: Shui, Bo, et al.
Pubblicazione: (2026)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
di: Gebreegziabher, Simret Araya, et al.
Pubblicazione: (2026)
di: Gebreegziabher, Simret Araya, et al.
Pubblicazione: (2026)
SPHERE: Scaling Personalized Feedback in Programming Classrooms with Structured Review of LLM Outputs
di: Tang, Xiaohang, et al.
Pubblicazione: (2024)
di: Tang, Xiaohang, et al.
Pubblicazione: (2024)
Building AI Literacy at Home: How Families Navigate Children's Self-Directed Learning with AI
di: Xie, Jingyi, et al.
Pubblicazione: (2025)
di: Xie, Jingyi, et al.
Pubblicazione: (2025)
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
di: Tan, Felicia Fang-Yi, et al.
Pubblicazione: (2026)
di: Tan, Felicia Fang-Yi, et al.
Pubblicazione: (2026)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
di: Chen, Chaoran, et al.
Pubblicazione: (2025)
di: Chen, Chaoran, et al.
Pubblicazione: (2025)
Adanonymizer: Interactively Navigating and Balancing the Duality of Privacy and Output Performance in Human-LLM Interaction
di: Zhang, Shuning, et al.
Pubblicazione: (2024)
di: Zhang, Shuning, et al.
Pubblicazione: (2024)
The Observability Gap: Why Output-Level Human Feedback Fails for LLM Coding Agents
di: Wang, Yinghao, et al.
Pubblicazione: (2026)
di: Wang, Yinghao, et al.
Pubblicazione: (2026)
Task-Aware Delegation Cues for LLM Agents
di: Gu, Xingrui
Pubblicazione: (2026)
di: Gu, Xingrui
Pubblicazione: (2026)
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
di: Nogueira, Brenda, et al.
Pubblicazione: (2025)
di: Nogueira, Brenda, et al.
Pubblicazione: (2025)
Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
di: Elangovan, Aparna, et al.
Pubblicazione: (2024)
di: Elangovan, Aparna, et al.
Pubblicazione: (2024)
KCluster: An LLM-based Clustering Approach to Knowledge Component Discovery
di: Wei, Yumou, et al.
Pubblicazione: (2025)
di: Wei, Yumou, et al.
Pubblicazione: (2025)
Designing an LLM-Based Behavioral Activation Chatbot for Young People with Depression: Insights from an Evaluation with Artificial Users and Clinical Experts
di: Kuhlmeier, Florian Onur, et al.
Pubblicazione: (2025)
di: Kuhlmeier, Florian Onur, et al.
Pubblicazione: (2025)
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
di: Yao, Bingsheng, et al.
Pubblicazione: (2026)
di: Yao, Bingsheng, et al.
Pubblicazione: (2026)
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
CPS-TaskForge: Generating Collaborative Problem Solving Environments for Diverse Communication Tasks
di: Haduong, Nikita, et al.
Pubblicazione: (2024)
di: Haduong, Nikita, et al.
Pubblicazione: (2024)
Exploring Expert Perspectives on Wearable-Triggered LLM Conversational Support for Daily Stress Management
di: Dongre, Poorvesh, et al.
Pubblicazione: (2026)
di: Dongre, Poorvesh, et al.
Pubblicazione: (2026)
DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users
di: Kawabe, Wataru, et al.
Pubblicazione: (2024)
di: Kawabe, Wataru, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
di: Szymanski, Annalisa, et al.
Pubblicazione: (2026) -
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
di: Szymanski, Annalisa, et al.
Pubblicazione: (2024) -
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
di: Chiang, Charles, et al.
Pubblicazione: (2026) -
"I'm categorizing LLM as a productivity tool": Examining ethics of LLM use in HCI research practices
di: Kapania, Shivani, et al.
Pubblicazione: (2024) -
Large Language Model Agent Personality and Response Appropriateness: Evaluation by Human Linguistic Experts, LLM-as-Judge, and Natural Language Processing Model
di: Jayakumar, Eswari, et al.
Pubblicazione: (2025)