The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Calderon, Nitay, Reichart, Roi, Dror, Rotem |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs
by: Calderon, Nitay, et al.
Published: (2024)
by: Calderon, Nitay, et al.
Published: (2024)
Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games
by: Shapira, Eilam, et al.
Published: (2024)
by: Shapira, Eilam, et al.
Published: (2024)
Can Unconfident LLM Annotations Be Used for Confident Conclusions?
by: Gligorić, Kristina, et al.
Published: (2024)
by: Gligorić, Kristina, et al.
Published: (2024)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
by: Li, Weiyue, et al.
Published: (2026)
by: Li, Weiyue, et al.
Published: (2026)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
by: Kim, Jane Paik
Published: (2026)
by: Kim, Jane Paik
Published: (2026)
Model-in-the-Loop (MILO): Accelerating Multimodal AI Data Annotation with LLMs
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation
by: Wenz, Fabian, et al.
Published: (2025)
by: Wenz, Fabian, et al.
Published: (2025)
Investigating Low-Cost LLM Annotation for~Spoken Dialogue Understanding Datasets
by: Druart, Lucas, et al.
Published: (2024)
by: Druart, Lucas, et al.
Published: (2024)
QACP: An Annotated Question Answering Dataset for Assisting Chinese Python Programming Learners
by: Xiao, Rui, et al.
Published: (2024)
by: Xiao, Rui, et al.
Published: (2024)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025)
by: McCullum, Lucas, et al.
Published: (2025)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
by: Toker, Gilat, et al.
Published: (2026)
by: Toker, Gilat, et al.
Published: (2026)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
If in a Crowdsourced Data Annotation Pipeline, a GPT-4
by: He, Zeyu, et al.
Published: (2024)
by: He, Zeyu, et al.
Published: (2024)
Fakes of Varying Shades: How Warning Affects Human Perception and Engagement Regarding LLM Hallucinations
by: Nahar, Mahjabin, et al.
Published: (2024)
by: Nahar, Mahjabin, et al.
Published: (2024)
MEGAnno+: A Human-LLM Collaborative Annotation System
by: Kim, Hannah, et al.
Published: (2024)
by: Kim, Hannah, et al.
Published: (2024)
Leveraging Large Language Models (LLMs) to Support Collaborative Human-AI Online Risk Data Annotation
by: Park, Jinkyung, et al.
Published: (2024)
by: Park, Jinkyung, et al.
Published: (2024)
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
by: Zhu, Xiaoyuan, et al.
Published: (2026)
by: Zhu, Xiaoyuan, et al.
Published: (2026)
Evaluating LLMs as Human Surrogates in Controlled Experiments
by: Hoq, Adnan, et al.
Published: (2026)
by: Hoq, Adnan, et al.
Published: (2026)
Text-to-SQL Domain Adaptation via Human-LLM Collaborative Data Annotation
by: Tian, Yuan, et al.
Published: (2025)
by: Tian, Yuan, et al.
Published: (2025)
Game Development as Human-LLM Interaction
by: Hong, Jiale, et al.
Published: (2024)
by: Hong, Jiale, et al.
Published: (2024)
HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?
by: Ji, Sijie, et al.
Published: (2024)
by: Ji, Sijie, et al.
Published: (2024)
Creative Beam Search: LLM-as-a-Judge For Improving Response Generation
by: Franceschelli, Giorgio, et al.
Published: (2024)
by: Franceschelli, Giorgio, et al.
Published: (2024)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
Retrieve, Annotate, Evaluate, Repeat: Leveraging Multimodal LLMs for Large-Scale Product Retrieval Evaluation
by: Hosseini, Kasra, et al.
Published: (2024)
by: Hosseini, Kasra, et al.
Published: (2024)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
by: Palta, Shramay, et al.
Published: (2025)
by: Palta, Shramay, et al.
Published: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
by: Lee, Dong Won, et al.
Published: (2024)
by: Lee, Dong Won, et al.
Published: (2024)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Can LLMs Assist Annotators in Identifying Morality Frames? -- Case Study on Vaccination Debate on Social Media
by: Islam, Tunazzina, et al.
Published: (2025)
by: Islam, Tunazzina, et al.
Published: (2025)
Integrating Personality into Digital Humans: A Review of LLM-Driven Approaches for Virtual Reality
by: Brito, Iago Alves, et al.
Published: (2025)
by: Brito, Iago Alves, et al.
Published: (2025)
The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance
by: Cohen, Ruth, et al.
Published: (2026)
by: Cohen, Ruth, et al.
Published: (2026)
A Piece of Theatre: Investigating How Teachers Design LLM Chatbots to Assist Adolescent Cyberbullying Education
by: Hedderich, Michael A., et al.
Published: (2024)
by: Hedderich, Michael A., et al.
Published: (2024)
How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond
by: Huang, Chen, et al.
Published: (2025)
by: Huang, Chen, et al.
Published: (2025)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
Similar Items
-
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024) -
On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs
by: Calderon, Nitay, et al.
Published: (2024) -
Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games
by: Shapira, Eilam, et al.
Published: (2024) -
Can Unconfident LLM Annotations Be Used for Confident Conclusions?
by: Gligorić, Kristina, et al.
Published: (2024) -
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
by: Li, Weiyue, et al.
Published: (2026)