Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Tzu-Heng, Vishwakarma, Harit, Sala, Frederic
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913890061778944
author Huang, Tzu-Heng
Vishwakarma, Harit
Sala, Frederic
author_facet Huang, Tzu-Heng
Vishwakarma, Harit
Sala, Frederic
contents Large language models (LLMs) are widely used to evaluate the quality of LLM generations and responses, but this leads to significant challenges: high API costs, uncertain reliability, inflexible pipelines, and inherent biases. To address these, we introduce PAJAMA (Program-As-a-Judge for Automated Model Assessment), a new alternative that uses LLMs to synthesize executable judging programs instead of directly scoring responses. These synthesized programs can be stored and run locally, costing orders of magnitude less while providing interpretable, and auditable judging logic that can be easily adapted. Program-based judges mitigate biases, improving judgment consistency by 15.83% and reducing biased responses by 23.7% on average compared to a Qwen2.5-14B-based LLM-as-a-judge. When program judgments are distilled into a model, PAJAMA outperforms LLM-as-a-judge on the challenging CHAT-HARD subset of RewardBench, outperforming metrics by 2.19% on Prometheus and 8.67% on the JudgeLM dataset, all at three orders of magnitude lower cost.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10403
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
Huang, Tzu-Heng
Vishwakarma, Harit
Sala, Frederic
Machine Learning
Artificial Intelligence
Large language models (LLMs) are widely used to evaluate the quality of LLM generations and responses, but this leads to significant challenges: high API costs, uncertain reliability, inflexible pipelines, and inherent biases. To address these, we introduce PAJAMA (Program-As-a-Judge for Automated Model Assessment), a new alternative that uses LLMs to synthesize executable judging programs instead of directly scoring responses. These synthesized programs can be stored and run locally, costing orders of magnitude less while providing interpretable, and auditable judging logic that can be easily adapted. Program-based judges mitigate biases, improving judgment consistency by 15.83% and reducing biased responses by 23.7% on average compared to a Qwen2.5-14B-based LLM-as-a-judge. When program judgments are distilled into a model, PAJAMA outperforms LLM-as-a-judge on the challenging CHAT-HARD subset of RewardBench, outperforming metrics by 2.19% on Prometheus and 8.67% on the JudgeLM dataset, all at three orders of magnitude lower cost.
title Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.10403