Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kim, Jane Paik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Creative Beam Search: LLM-as-a-Judge For Improving Response Generation
von: Franceschelli, Giorgio, et al.
Veröffentlicht: (2024)
von: Franceschelli, Giorgio, et al.
Veröffentlicht: (2024)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
von: Sturgeon, Benjamin, et al.
Veröffentlicht: (2025)
von: Sturgeon, Benjamin, et al.
Veröffentlicht: (2025)
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
von: Xi, Sarina, et al.
Veröffentlicht: (2025)
von: Xi, Sarina, et al.
Veröffentlicht: (2025)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
von: Soni, Nikita, et al.
Veröffentlicht: (2025)
von: Soni, Nikita, et al.
Veröffentlicht: (2025)
Value Profiles for Encoding Human Variation
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
Addressing the Ecological Fallacy in Larger LMs with Human Context
von: Soni, Nikita, et al.
Veröffentlicht: (2026)
von: Soni, Nikita, et al.
Veröffentlicht: (2026)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
von: Abels, Axel, et al.
Veröffentlicht: (2025)
von: Abels, Axel, et al.
Veröffentlicht: (2025)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
von: Li, Weiyue, et al.
Veröffentlicht: (2026)
von: Li, Weiyue, et al.
Veröffentlicht: (2026)
Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual Relationships
von: Boggust, Angie, et al.
Veröffentlicht: (2024)
von: Boggust, Angie, et al.
Veröffentlicht: (2024)
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
von: Saha, Shoumik, et al.
Veröffentlicht: (2025)
von: Saha, Shoumik, et al.
Veröffentlicht: (2025)
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems
von: Rajagopal, Shreya, et al.
Veröffentlicht: (2024)
von: Rajagopal, Shreya, et al.
Veröffentlicht: (2024)
DiscoverLLM: From Executing Intents to Discovering Them
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026)
HybridQuestion: Human-AI Collaboration for Identifying High-Impact Research Questions
von: Zhao, Keyu, et al.
Veröffentlicht: (2025)
von: Zhao, Keyu, et al.
Veröffentlicht: (2025)
Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
von: Jain, Raunak
Veröffentlicht: (2025)
von: Jain, Raunak
Veröffentlicht: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
von: Xu, Ziwen, et al.
Veröffentlicht: (2026)
von: Xu, Ziwen, et al.
Veröffentlicht: (2026)
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
von: Kahng, Minsuk, et al.
Veröffentlicht: (2024)
von: Kahng, Minsuk, et al.
Veröffentlicht: (2024)
Align When They Want, Complement When They Need! Human-Centered Ensembles for Adaptive Human-AI Collaboration
von: Amin, Hasan, et al.
Veröffentlicht: (2026)
von: Amin, Hasan, et al.
Veröffentlicht: (2026)
Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
von: Zhang, Ziyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyuan, et al.
Veröffentlicht: (2025)
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
von: He, Zeyu, et al.
Veröffentlicht: (2025)
von: He, Zeyu, et al.
Veröffentlicht: (2025)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
von: Park, Jung In, et al.
Veröffentlicht: (2024)
von: Park, Jung In, et al.
Veröffentlicht: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
von: Baidya, Avinash, et al.
Veröffentlicht: (2025)
von: Baidya, Avinash, et al.
Veröffentlicht: (2025)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
von: Sung, Yoo Yeon, et al.
Veröffentlicht: (2025)
von: Sung, Yoo Yeon, et al.
Veröffentlicht: (2025)
UniAutoML: A Human-Centered Framework for Unified Discriminative and Generative AutoML with Large Language Models
von: Guo, Jiayi, et al.
Veröffentlicht: (2024)
von: Guo, Jiayi, et al.
Veröffentlicht: (2024)
Are You Being Tracked? Discover the Power of Zero-Shot Trajectory Tracing with LLMs!
von: Yang, Huanqi, et al.
Veröffentlicht: (2024)
von: Yang, Huanqi, et al.
Veröffentlicht: (2024)
LLM Attributor: Interactive Visual Attribution for LLM Generation
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration
von: Grannen, Jennifer, et al.
Veröffentlicht: (2024)
von: Grannen, Jennifer, et al.
Veröffentlicht: (2024)
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
von: Grannen, Jennifer, et al.
Veröffentlicht: (2025)
von: Grannen, Jennifer, et al.
Veröffentlicht: (2025)
Automating Customer Needs Analysis: A Comparative Study of Large Language Models in the Travel Industry
von: Barandoni, Simone, et al.
Veröffentlicht: (2024)
von: Barandoni, Simone, et al.
Veröffentlicht: (2024)
Human-Computer Interaction and Human-AI Collaboration in Advanced Air Mobility: A Comprehensive Review
von: Sagirli, Fatma Yamac, et al.
Veröffentlicht: (2024)
von: Sagirli, Fatma Yamac, et al.
Veröffentlicht: (2024)
Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task Automation
von: Lee, Sunjae, et al.
Veröffentlicht: (2023)
von: Lee, Sunjae, et al.
Veröffentlicht: (2023)
AI Agents for Inventory Control: Human-LLM-OR Complementarity
von: Baek, Jackie, et al.
Veröffentlicht: (2026)
von: Baek, Jackie, et al.
Veröffentlicht: (2026)
Augmenting Automation: Intent-Based User Instruction Classification with Machine Learning
von: Basyal, Lochan, et al.
Veröffentlicht: (2024)
von: Basyal, Lochan, et al.
Veröffentlicht: (2024)
KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents
von: Zhu, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2024)
Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
von: Liu, Jonathan, et al.
Veröffentlicht: (2025)
von: Liu, Jonathan, et al.
Veröffentlicht: (2025)
Cognitive Exoskeleton: Augmenting Human Cognition with an AI-Mediated Intelligent Visual Feedback
von: Xu, Songlin, et al.
Veröffentlicht: (2025)
von: Xu, Songlin, et al.
Veröffentlicht: (2025)
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
von: Seshadri, Preethi, et al.
Veröffentlicht: (2026)
von: Seshadri, Preethi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Creative Beam Search: LLM-as-a-Judge For Improving Response Generation
von: Franceschelli, Giorgio, et al.
Veröffentlicht: (2024) -
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025) -
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
von: Sturgeon, Benjamin, et al.
Veröffentlicht: (2025) -
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
von: Xi, Sarina, et al.
Veröffentlicht: (2025) -
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
von: Soni, Nikita, et al.
Veröffentlicht: (2025)