Verdict: A Library for Scaling Judge-Time Compute
Fuente:
arXiv
Guardado en:
| Autores principales: | Kalra, Nimit, Tang, Leonard |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
por: Han, Steve, et al.
Publicado: (2025)
por: Han, Steve, et al.
Publicado: (2025)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
por: Badshah, Sher, et al.
Publicado: (2024)
por: Badshah, Sher, et al.
Publicado: (2024)
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
por: Kong, Chuyi, et al.
Publicado: (2025)
por: Kong, Chuyi, et al.
Publicado: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
por: Zhou, Yilun, et al.
Publicado: (2025)
por: Zhou, Yilun, et al.
Publicado: (2025)
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
por: Liu, Yuhan, et al.
Publicado: (2025)
por: Liu, Yuhan, et al.
Publicado: (2025)
Domain Adaptation Through Task Distillation
por: Zhou, Brady, et al.
Publicado: (2020)
por: Zhou, Brady, et al.
Publicado: (2020)
Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean
por: Meshkov, Aleksandr
Publicado: (2026)
por: Meshkov, Aleksandr
Publicado: (2026)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
por: Chan, Chi-Min, et al.
Publicado: (2025)
por: Chan, Chi-Min, et al.
Publicado: (2025)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
por: Tang, Zhenwei, et al.
Publicado: (2026)
por: Tang, Zhenwei, et al.
Publicado: (2026)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
por: Bi, Zhenni, et al.
Publicado: (2024)
por: Bi, Zhenni, et al.
Publicado: (2024)
LePaRD: A Large-Scale Dataset of Judges Citing Precedents
por: Mahari, Robert, et al.
Publicado: (2023)
por: Mahari, Robert, et al.
Publicado: (2023)
Inverse Scaling in Test-Time Compute
por: Gema, Aryo Pradipta, et al.
Publicado: (2025)
por: Gema, Aryo Pradipta, et al.
Publicado: (2025)
The Art of Scaling Test-Time Compute for Large Language Models
por: Agarwal, Aradhye, et al.
Publicado: (2025)
por: Agarwal, Aradhye, et al.
Publicado: (2025)
Chain of Methodologies: Scaling Test Time Computation without Training
por: Liu, Cong, et al.
Publicado: (2025)
por: Liu, Cong, et al.
Publicado: (2025)
Parallel Loop Transformer for Efficient Test-Time Computation Scaling
por: Wu, Bohong, et al.
Publicado: (2025)
por: Wu, Bohong, et al.
Publicado: (2025)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
por: Liu, Shuliang, et al.
Publicado: (2025)
por: Liu, Shuliang, et al.
Publicado: (2025)
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
por: Ho, Xanh, et al.
Publicado: (2025)
por: Ho, Xanh, et al.
Publicado: (2025)
MR. Judge: Multimodal Reasoner as a Judge
por: Pi, Renjie, et al.
Publicado: (2025)
por: Pi, Renjie, et al.
Publicado: (2025)
LLM-as-a-Judge for Time Series Explanations
por: Sivalingam, Preetham, et al.
Publicado: (2026)
por: Sivalingam, Preetham, et al.
Publicado: (2026)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
por: Le, Hieu Xuan, et al.
Publicado: (2026)
por: Le, Hieu Xuan, et al.
Publicado: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
por: Belmadani, Ikram, et al.
Publicado: (2026)
por: Belmadani, Ikram, et al.
Publicado: (2026)
Adaptive Rectification Sampling for Test-Time Compute Scaling
por: Tan, Zhendong, et al.
Publicado: (2025)
por: Tan, Zhendong, et al.
Publicado: (2025)
Test-Time Scaling Makes Overtraining Compute-Optimal
por: Roberts, Nicholas, et al.
Publicado: (2026)
por: Roberts, Nicholas, et al.
Publicado: (2026)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
por: Shi, Lin, et al.
Publicado: (2024)
por: Shi, Lin, et al.
Publicado: (2024)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
por: Bellibatlu, Rohith Reddy, et al.
Publicado: (2026)
por: Bellibatlu, Rohith Reddy, et al.
Publicado: (2026)
Computer-Use Agents as Judges for Generative User Interface
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
por: Lin, Kevin Qinghong, et al.
Publicado: (2025)
Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
por: Ning, Meiling, et al.
Publicado: (2025)
por: Ning, Meiling, et al.
Publicado: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
por: Thakur, Aman Singh, et al.
Publicado: (2024)
por: Thakur, Aman Singh, et al.
Publicado: (2024)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
por: Marioriyad, Arash, et al.
Publicado: (2025)
por: Marioriyad, Arash, et al.
Publicado: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
por: Yang, Bo, et al.
Publicado: (2026)
por: Yang, Bo, et al.
Publicado: (2026)
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
por: Bavaresco, Anna, et al.
Publicado: (2024)
por: Bavaresco, Anna, et al.
Publicado: (2024)
Scaling Natural-Language Graph-Based Test Time Compute for Automated Theorem Proving
por: Li, Vincent, et al.
Publicado: (2025)
por: Li, Vincent, et al.
Publicado: (2025)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
por: Setlur, Amrith, et al.
Publicado: (2025)
por: Setlur, Amrith, et al.
Publicado: (2025)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
por: Zhu, Ziyi, et al.
Publicado: (2026)
por: Zhu, Ziyi, et al.
Publicado: (2026)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
por: Wang, Leyao, et al.
Publicado: (2026)
por: Wang, Leyao, et al.
Publicado: (2026)
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
por: DeLucia, Alexandra, et al.
Publicado: (2026)
por: DeLucia, Alexandra, et al.
Publicado: (2026)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
por: Wang, Qian, et al.
Publicado: (2025)
por: Wang, Qian, et al.
Publicado: (2025)
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
por: Zhao, Jian, et al.
Publicado: (2025)
por: Zhao, Jian, et al.
Publicado: (2025)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
por: Yang, Wenkai, et al.
Publicado: (2025)
por: Yang, Wenkai, et al.
Publicado: (2025)
Ejemplares similares
-
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
por: Han, Steve, et al.
Publicado: (2025) -
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
por: Badshah, Sher, et al.
Publicado: (2024) -
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
por: Kong, Chuyi, et al.
Publicado: (2025) -
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
por: Zhou, Yilun, et al.
Publicado: (2025) -
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
por: Liu, Yuhan, et al.
Publicado: (2025)