Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Verga, Pat, Hofstatter, Sebastian, Althammer, Sophia, Su, Yixuan, Piktus, Aleksandra, Arkhangorodsky, Arkady, Xu, Minjie, White, Naomi, Lewis, Patrick |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
di: Li, Xiaochuan, et al.
Pubblicazione: (2025)
di: Li, Xiaochuan, et al.
Pubblicazione: (2025)
ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering
di: Li, Huayang, et al.
Pubblicazione: (2024)
di: Li, Huayang, et al.
Pubblicazione: (2024)
FLARE: Faithful Logic-Aided Reasoning and Exploration
di: Arakelyan, Erik, et al.
Pubblicazione: (2024)
di: Arakelyan, Erik, et al.
Pubblicazione: (2024)
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
di: Reese, May Lynn, et al.
Pubblicazione: (2026)
di: Reese, May Lynn, et al.
Pubblicazione: (2026)
Comparing Human and LLM Generated Code: The Jury is Still Out!
di: Licorish, Sherlock A., et al.
Pubblicazione: (2025)
di: Licorish, Sherlock A., et al.
Pubblicazione: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
di: Wang, Ruiqi, et al.
Pubblicazione: (2025)
di: Wang, Ruiqi, et al.
Pubblicazione: (2025)
A Finite-Calibration Regime Map for LLM Judge Panels
di: Zhu, Bin, et al.
Pubblicazione: (2026)
di: Zhu, Bin, et al.
Pubblicazione: (2026)
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
di: Shi, Sherry, et al.
Pubblicazione: (2025)
di: Shi, Sherry, et al.
Pubblicazione: (2025)
Jury: A Comprehensive Evaluation Toolkit
di: Cavusoglu, Devrim, et al.
Pubblicazione: (2023)
di: Cavusoglu, Devrim, et al.
Pubblicazione: (2023)
Juris
Pubblicazione: (2017)
Pubblicazione: (2017)
Evaluating the Well-Being of Public Library Workers
di: Juniper, Bridget, et al.
Pubblicazione: (2012)
di: Juniper, Bridget, et al.
Pubblicazione: (2012)
Definição e Percepção de Imagem: Um Estudo em uma Escola de Educação Infantil de Novo Hamburgo
di: Cássia Rebelo Hofstätter
Pubblicazione: (2008)
di: Cássia Rebelo Hofstätter
Pubblicazione: (2008)
A Pesquisa de Marketing como um Meio de Informação para a Tomada de Decisão Estratégica
di: Cássia Rebelo Hofstätter
Pubblicazione: (2005)
di: Cássia Rebelo Hofstätter
Pubblicazione: (2005)
Imagem e Identidade Institucional: Um Estudo Aplicado à Feevale
di: Cássia Rebello Hofstätter
Pubblicazione: (2009)
di: Cássia Rebello Hofstätter
Pubblicazione: (2009)
Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity
di: Engländer, Leon, et al.
Pubblicazione: (2026)
di: Engländer, Leon, et al.
Pubblicazione: (2026)
Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
di: Wu, Xuanxin, et al.
Pubblicazione: (2025)
di: Wu, Xuanxin, et al.
Pubblicazione: (2025)
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face
di: Akiki, Christopher, et al.
Pubblicazione: (2023)
di: Akiki, Christopher, et al.
Pubblicazione: (2023)
Chapter Lo Studio, le accademie, la città (secc. XVI-XIX)
di: Verga, Marcello
Pubblicazione: (2024)
di: Verga, Marcello
Pubblicazione: (2024)
Rodales semilleros de Prosopis a partir del bosque nativo
di: A. Verga
Pubblicazione: (2014)
di: A. Verga
Pubblicazione: (2014)
O bem-estar subjetivo no comportamento de compra de alimentos orgânicos
di: Everton Verga
Pubblicazione: (2020)
di: Everton Verga
Pubblicazione: (2020)
Atitudes maternas face à amamentação e satisfação com o suporte social
di: Vanessa Verga
Pubblicazione: (2022)
di: Vanessa Verga
Pubblicazione: (2022)
EMPREENDEDORISMO: EVOLUÇÃO HISTÓRICA, DEFINIÇÕES E ABORDAGENS.
di: Everton Verga
Pubblicazione: (2014)
di: Everton Verga
Pubblicazione: (2014)
CABALLERÍA RUSTICANA
di: Giovanni Verga
Pubblicazione: (2009)
di: Giovanni Verga
Pubblicazione: (2009)
Caracterización morfológica de los algarrobos (Prosopis sp.) en las regiones fitogeográficas Chaqueña y Espinal norte de Argentina
di: A. Verga
Pubblicazione: (2009)
di: A. Verga
Pubblicazione: (2009)
Vox Juris
Pubblicazione: (2018)
Pubblicazione: (2018)
Ratio Juris
Pubblicazione: (2018)
Pubblicazione: (2018)
Jury Trial
di: Igor Yurievich NIKODIMOV
Pubblicazione: (2020)
di: Igor Yurievich NIKODIMOV
Pubblicazione: (2020)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
di: Wei, Hui, et al.
Pubblicazione: (2024)
di: Wei, Hui, et al.
Pubblicazione: (2024)
Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress
di: Odumakinde, Ayomide, et al.
Pubblicazione: (2024)
di: Odumakinde, Ayomide, et al.
Pubblicazione: (2024)
Generalizing Condorcet's Jury Theorem to Social Networks
di: Braha, Dan, et al.
Pubblicazione: (2025)
di: Braha, Dan, et al.
Pubblicazione: (2025)
Uma contribuição da educação ambiental crítica para (des)construção do olhar sobre a seca no semiárido baiano
di: Lakshmi Juliane Vallim Hofstatter
Pubblicazione: (2016)
di: Lakshmi Juliane Vallim Hofstatter
Pubblicazione: (2016)
Ancient asexuality: No scandals found with novel data
di: Paulo Hofstatter, et al.
Pubblicazione: (2024)
di: Paulo Hofstatter, et al.
Pubblicazione: (2024)
Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries
di: Ullah, Muhammad Aziz, et al.
Pubblicazione: (2026)
di: Ullah, Muhammad Aziz, et al.
Pubblicazione: (2026)
School Counselors and Teacher-Librarians: A Necessary Partnership for Effective Schools.
di: White, Maureen, et al.
Pubblicazione: (1997)
di: White, Maureen, et al.
Pubblicazione: (1997)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
di: Chen, Junjie, et al.
Pubblicazione: (2026)
di: Chen, Junjie, et al.
Pubblicazione: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
Evaluating Metrics for Safety with LLM-as-Judges
di: Clegg, Kester, et al.
Pubblicazione: (2025)
di: Clegg, Kester, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
di: Li, Xiaochuan, et al.
Pubblicazione: (2025) -
ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering
di: Li, Huayang, et al.
Pubblicazione: (2024) -
FLARE: Faithful Logic-Aided Reasoning and Exploration
di: Arakelyan, Erik, et al.
Pubblicazione: (2024) -
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
di: Reese, May Lynn, et al.
Pubblicazione: (2026) -
Comparing Human and LLM Generated Code: The Jury is Still Out!
di: Licorish, Sherlock A., et al.
Pubblicazione: (2025)