Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Tzu-Heng, Vishwakarma, Harit, Sala, Frederic |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Promises and Pitfalls of Threshold-based Auto-labeling
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2022)
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2022)
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
von: Shin, Changho, et al.
Veröffentlicht: (2024)
von: Shin, Changho, et al.
Veröffentlicht: (2024)
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
von: Zhao, Jitian, et al.
Veröffentlicht: (2026)
von: Zhao, Jitian, et al.
Veröffentlicht: (2026)
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
von: Bauer, Justin, et al.
Veröffentlicht: (2026)
von: Bauer, Justin, et al.
Veröffentlicht: (2026)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2024)
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2024)
Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
von: Yamada, Daisuke, et al.
Veröffentlicht: (2025)
von: Yamada, Daisuke, et al.
Veröffentlicht: (2025)
Taming False Positives in Out-of-Distribution Detection with Human Feedback
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2024)
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2024)
ScriptoriumWS: A Code Generation Assistant for Weak Supervision
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2024)
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2024)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2026)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2026)
MoRe Fine-Tuning with 10x Fewer Parameters
von: Tan, Wenxuan, et al.
Veröffentlicht: (2024)
von: Tan, Wenxuan, et al.
Veröffentlicht: (2024)
Causal Spherical Hypergraph Networks for Modelling Social Uncertainty
von: Harit, Anoushka, et al.
Veröffentlicht: (2025)
von: Harit, Anoushka, et al.
Veröffentlicht: (2025)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
Actionable Interpretability via Causal Hypergraphs: Unravelling Batch Size Effects in Deep Learning
von: Sun, Zhongtian, et al.
Veröffentlicht: (2025)
von: Sun, Zhongtian, et al.
Veröffentlicht: (2025)
RicciFlowRec: A Geometric Root Cause Recommender Using Ricci Curvature on Financial Graphs
von: Sun, Zhongtian, et al.
Veröffentlicht: (2025)
von: Sun, Zhongtian, et al.
Veröffentlicht: (2025)
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
von: Vishwakarma, Harsh, et al.
Veröffentlicht: (2025)
von: Vishwakarma, Harsh, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
von: Halder, Indranil, et al.
Veröffentlicht: (2025)
von: Halder, Indranil, et al.
Veröffentlicht: (2025)
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
Weak-to-Strong Generalization Through the Data-Centric Lens
von: Shin, Changho, et al.
Veröffentlicht: (2024)
von: Shin, Changho, et al.
Veröffentlicht: (2024)
Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
von: Fabbri, Francesco, et al.
Veröffentlicht: (2025)
von: Fabbri, Francesco, et al.
Veröffentlicht: (2025)
Auto-Prompt Ensemble for LLM Judge
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
ManifoldMind: Dynamic Hyperbolic Reasoning for Trustworthy Recommendations
von: Harit, Anoushka, et al.
Veröffentlicht: (2025)
von: Harit, Anoushka, et al.
Veröffentlicht: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Breaking Down Financial News Impact: A Novel AI Approach with Geometric Hypergraphs
von: Harit, Anoushka, et al.
Veröffentlicht: (2024)
von: Harit, Anoushka, et al.
Veröffentlicht: (2024)
Optimal partition of feature using Bayesian classifier
von: Vishwakarma, Sanjay, et al.
Veröffentlicht: (2023)
von: Vishwakarma, Sanjay, et al.
Veröffentlicht: (2023)
Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation
von: Zhao, Jitian, et al.
Veröffentlicht: (2025)
von: Zhao, Jitian, et al.
Veröffentlicht: (2025)
Zero-Shot Robustification of Zero-Shot Models
von: Adila, Dyah, et al.
Veröffentlicht: (2023)
von: Adila, Dyah, et al.
Veröffentlicht: (2023)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
von: Wagner, Nico, et al.
Veröffentlicht: (2024)
von: Wagner, Nico, et al.
Veröffentlicht: (2024)
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
von: Zhang, Yasi, et al.
Veröffentlicht: (2026)
von: Zhang, Yasi, et al.
Veröffentlicht: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
von: Ge, Albert, et al.
Veröffentlicht: (2025)
von: Ge, Albert, et al.
Veröffentlicht: (2025)
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
von: Singh, Janvijay, et al.
Veröffentlicht: (2025)
von: Singh, Janvijay, et al.
Veröffentlicht: (2025)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
von: Dadkhahi, Hamid, et al.
Veröffentlicht: (2025)
von: Dadkhahi, Hamid, et al.
Veröffentlicht: (2025)
Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
von: Jwa, Seungyeon, et al.
Veröffentlicht: (2025)
von: Jwa, Seungyeon, et al.
Veröffentlicht: (2025)
GLANCE: Graph Logic Attention Network with Cluster Enhancement for Heterophilous Graph Representation Learning
von: Sun, Zhongtian, et al.
Veröffentlicht: (2025)
von: Sun, Zhongtian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Promises and Pitfalls of Threshold-based Auto-labeling
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2022) -
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
von: Shin, Changho, et al.
Veröffentlicht: (2024) -
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
von: Zhao, Jitian, et al.
Veröffentlicht: (2026) -
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
von: Bauer, Justin, et al.
Veröffentlicht: (2026) -
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)