Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Findeis, Arduin, Weers, Floris, Yin, Guoli, Ye, Ke, Pang, Ruoming, Gunter, Tom |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Feedback Forensics: A Toolkit to Measure AI Personality
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
Inverse Constitutional AI: Compressing Preferences into Principles
von: Findeis, Arduin, et al.
Veröffentlicht: (2024)
von: Findeis, Arduin, et al.
Veröffentlicht: (2024)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
von: Lu, Jiarui, et al.
Veröffentlicht: (2024)
von: Lu, Jiarui, et al.
Veröffentlicht: (2024)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Distillation Scaling Laws
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
von: Han, Steve, et al.
Veröffentlicht: (2025)
von: Han, Steve, et al.
Veröffentlicht: (2025)
Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce
von: Chen, Liang, et al.
Veröffentlicht: (2026)
von: Chen, Liang, et al.
Veröffentlicht: (2026)
Large Language Model-guided Document Selection
von: Kong, Xiang, et al.
Veröffentlicht: (2024)
von: Kong, Xiang, et al.
Veröffentlicht: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
von: Parkar, Ritik Sachin, et al.
Veröffentlicht: (2024)
von: Parkar, Ritik Sachin, et al.
Veröffentlicht: (2024)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
Can Unconfident LLM Annotations Be Used for Confident Conclusions?
von: Gligorić, Kristina, et al.
Veröffentlicht: (2024)
von: Gligorić, Kristina, et al.
Veröffentlicht: (2024)
X-AMR Annotation Tool
von: Ahmed, Shafiuddin Rehan, et al.
Veröffentlicht: (2024)
von: Ahmed, Shafiuddin Rehan, et al.
Veröffentlicht: (2024)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
ToolBridge: An Open-Source Dataset to Equip LLMs with External Tool Capabilities
von: Jin, Zhenchao, et al.
Veröffentlicht: (2024)
von: Jin, Zhenchao, et al.
Veröffentlicht: (2024)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding?
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
von: Amin, Hasan, et al.
Veröffentlicht: (2026)
von: Amin, Hasan, et al.
Veröffentlicht: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use
von: Pang, Renning, et al.
Veröffentlicht: (2026)
von: Pang, Renning, et al.
Veröffentlicht: (2026)
Synthetic bootstrapped pretraining
von: Yang, Zitong, et al.
Veröffentlicht: (2025)
von: Yang, Zitong, et al.
Veröffentlicht: (2025)
Can LLMs be Good Graph Judge for Knowledge Graph Construction?
von: Huang, Haoyu, et al.
Veröffentlicht: (2024)
von: Huang, Haoyu, et al.
Veröffentlicht: (2024)
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
von: Reese, May Lynn, et al.
Veröffentlicht: (2026)
von: Reese, May Lynn, et al.
Veröffentlicht: (2026)
Societal Alignment Frameworks Can Improve LLM Alignment
von: Stańczak, Karolina, et al.
Veröffentlicht: (2025)
von: Stańczak, Karolina, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Time Series Explanations
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs
von: Li, Yijie, et al.
Veröffentlicht: (2024)
von: Li, Yijie, et al.
Veröffentlicht: (2024)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
von: Li, Jidong, et al.
Veröffentlicht: (2025)
von: Li, Jidong, et al.
Veröffentlicht: (2025)
Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models
von: Liu, Yule, et al.
Veröffentlicht: (2025)
von: Liu, Yule, et al.
Veröffentlicht: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
Are We on the Right Way to Assessing LLM-as-a-Judge?
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Feedback Forensics: A Toolkit to Measure AI Personality
von: Findeis, Arduin, et al.
Veröffentlicht: (2025) -
Inverse Constitutional AI: Compressing Preferences into Principles
von: Findeis, Arduin, et al.
Veröffentlicht: (2024) -
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
von: Lu, Jiarui, et al.
Veröffentlicht: (2024) -
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025) -
Distillation Scaling Laws
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)