Enregistré dans:
| Auteur principal: | Zhang, Xinran |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2603.28005 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
par: Lee, Dongryeol, et autres
Publié: (2026)
par: Lee, Dongryeol, et autres
Publié: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
par: Belmadani, Ikram, et autres
Publié: (2026)
par: Belmadani, Ikram, et autres
Publié: (2026)
How Sensitive Are Safety Benchmarks to Judge Configuration Choices?
par: Zhang, Xinran
Publié: (2026)
par: Zhang, Xinran
Publié: (2026)
Cross-Lingual LLM-Judge Transfer via Evaluation Decomposition
par: Sheth, Ivaxi, et autres
Publié: (2026)
par: Sheth, Ivaxi, et autres
Publié: (2026)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
par: Bellibatlu, Rohith Reddy, et autres
Publié: (2026)
par: Bellibatlu, Rohith Reddy, et autres
Publié: (2026)
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
par: Adib, Shefayat E Shams, et autres
Publié: (2026)
par: Adib, Shefayat E Shams, et autres
Publié: (2026)
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
par: Elganayni, Mohamed Hesham, et autres
Publié: (2026)
par: Elganayni, Mohamed Hesham, et autres
Publié: (2026)
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
par: Ho, Xanh, et autres
Publié: (2025)
par: Ho, Xanh, et autres
Publié: (2025)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
par: Wei, Hui, et autres
Publié: (2024)
par: Wei, Hui, et autres
Publié: (2024)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
par: Badshah, Sher, et autres
Publié: (2024)
par: Badshah, Sher, et autres
Publié: (2024)
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
par: Hua, Yilun, et autres
Publié: (2026)
par: Hua, Yilun, et autres
Publié: (2026)
Structured Chain-of-Thought Prompting for Few-Shot Generation of Content-Grounded QA Conversations
par: Sultan, Md Arafat, et autres
Publié: (2024)
par: Sultan, Md Arafat, et autres
Publié: (2024)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
par: Zhang, Wentao, et autres
Publié: (2025)
par: Zhang, Wentao, et autres
Publié: (2025)
Prompting a Weighting Mechanism into LLM-as-a-Judge in Two-Step: A Case Study
par: Xie, Wenwen, et autres
Publié: (2025)
par: Xie, Wenwen, et autres
Publié: (2025)
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
par: Song, Mingyang, et autres
Publié: (2026)
par: Song, Mingyang, et autres
Publié: (2026)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
par: Lin, Wei-Hsiang, et autres
Publié: (2025)
par: Lin, Wei-Hsiang, et autres
Publié: (2025)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
par: Le, Hieu Xuan, et autres
Publié: (2026)
par: Le, Hieu Xuan, et autres
Publié: (2026)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
par: Huang, Hui, et autres
Publié: (2024)
par: Huang, Hui, et autres
Publié: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
par: Hong, Yihan, et autres
Publié: (2026)
par: Hong, Yihan, et autres
Publié: (2026)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
par: Shi, Lin, et autres
Publié: (2024)
par: Shi, Lin, et autres
Publié: (2024)
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
par: Zhang, Qiyuan, et autres
Publié: (2024)
par: Zhang, Qiyuan, et autres
Publié: (2024)
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
par: Gisserot-Boukhlef, Hippolyte, et autres
Publié: (2026)
par: Gisserot-Boukhlef, Hippolyte, et autres
Publié: (2026)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
par: Zhu, Ziyi, et autres
Publié: (2026)
par: Zhu, Ziyi, et autres
Publié: (2026)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
par: Krumdick, Michael, et autres
Publié: (2025)
par: Krumdick, Michael, et autres
Publié: (2025)
Same Content, Different Representations: A Controlled Study for Table QA
par: Zhang, Yue, et autres
Publié: (2025)
par: Zhang, Yue, et autres
Publié: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
par: Zhou, Yilun, et autres
Publié: (2025)
par: Zhou, Yilun, et autres
Publié: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
par: Tan, Sijun, et autres
Publié: (2024)
par: Tan, Sijun, et autres
Publié: (2024)
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
par: Maloyan, Narek, et autres
Publié: (2025)
par: Maloyan, Narek, et autres
Publié: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
par: Li, Qingquan, et autres
Publié: (2025)
par: Li, Qingquan, et autres
Publié: (2025)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
par: Li, Zhuochun, et autres
Publié: (2026)
par: Li, Zhuochun, et autres
Publié: (2026)
Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory
par: Sun, Jingwei, et autres
Publié: (2026)
par: Sun, Jingwei, et autres
Publié: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
par: Clegg, Kester, et autres
Publié: (2025)
par: Clegg, Kester, et autres
Publié: (2025)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
par: Lai, Viet Dac, et autres
Publié: (2024)
par: Lai, Viet Dac, et autres
Publié: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
par: Chen, Junjie, et autres
Publié: (2026)
par: Chen, Junjie, et autres
Publié: (2026)
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
par: Jonker, Richard A. A., et autres
Publié: (2026)
par: Jonker, Richard A. A., et autres
Publié: (2026)
JELV: A Judge of Edit-Level Validity for Evaluation and Automated Reference Expansion in Grammatical Error Correction
par: Zhan, Yuhao, et autres
Publié: (2025)
par: Zhan, Yuhao, et autres
Publié: (2025)
Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking
par: Zhang, Yichi, et autres
Publié: (2024)
par: Zhang, Yichi, et autres
Publié: (2024)
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
par: Yamauchi, Yusuke, et autres
Publié: (2025)
par: Yamauchi, Yusuke, et autres
Publié: (2025)
sebis at ArchEHR-QA 2026: How Much Can You Do Locally? Evaluating Grounded EHR QA on a Single Notebook
par: Yurt, Ibrahim Ebrar, et autres
Publié: (2026)
par: Yurt, Ibrahim Ebrar, et autres
Publié: (2026)
Neural at ArchEHR-QA 2025: Agentic Prompt Optimization for Evidence-Grounded Clinical Question Answering
par: Bogireddy, Sai Prasanna Teja Reddy, et autres
Publié: (2025)
par: Bogireddy, Sai Prasanna Teja Reddy, et autres
Publié: (2025)
Documents similaires
-
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
par: Lee, Dongryeol, et autres
Publié: (2026) -
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
par: Belmadani, Ikram, et autres
Publié: (2026) -
How Sensitive Are Safety Benchmarks to Judge Configuration Choices?
par: Zhang, Xinran
Publié: (2026) -
Cross-Lingual LLM-Judge Transfer via Evaluation Decomposition
par: Sheth, Ivaxi, et autres
Publié: (2026) -
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
par: Bellibatlu, Rohith Reddy, et autres
Publié: (2026)