RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Zhenwei, Liu, Zhaoyan, Hosseinzadeh, Rasa, Wu, Tongzi, Golestan, Keyvan, Cresswell, Jesse C. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DRESS: Disentangled Representation-based Self-Supervised Meta-Learning for Diverse Tasks
by: Cui, Wei, et al.
Published: (2025)
by: Cui, Wei, et al.
Published: (2025)
MSc-SQL: Multi-Sample Critiquing Small Language Models For Text-To-SQL Translation
by: Gorti, Satya Krishna, et al.
Published: (2024)
by: Gorti, Satya Krishna, et al.
Published: (2024)
Tabular Data Contrastive Learning via Class-Conditioned and Feature-Correlation Based Augmentation
by: Cui, Wei, et al.
Published: (2024)
by: Cui, Wei, et al.
Published: (2024)
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
by: Tang, Yuqi, et al.
Published: (2025)
by: Tang, Yuqi, et al.
Published: (2025)
A Geometric Framework for Understanding Memorization in Generative Models
by: Ross, Brendan Leigh, et al.
Published: (2024)
by: Ross, Brendan Leigh, et al.
Published: (2024)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
ObjexMT: Objective Extraction and Metacognitive Calibration for LLM-as-a-Judge under Multi-Turn Jailbreaks
by: Kim, Hyunjun, et al.
Published: (2025)
by: Kim, Hyunjun, et al.
Published: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
by: Chen, Junjie, et al.
Published: (2026)
by: Chen, Junjie, et al.
Published: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
by: Belmadani, Ikram, et al.
Published: (2026)
by: Belmadani, Ikram, et al.
Published: (2026)
Attribution Quality in AI-Generated Content:Benchmarking Style Embeddings and LLM Judges
by: Abbas, Misam
Published: (2025)
by: Abbas, Misam
Published: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
by: Chen, Jiamin, et al.
Published: (2026)
by: Chen, Jiamin, et al.
Published: (2026)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
by: Huang, Hui, et al.
Published: (2024)
by: Huang, Hui, et al.
Published: (2024)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
by: Zhu, Ziyi, et al.
Published: (2026)
by: Zhu, Ziyi, et al.
Published: (2026)
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
by: Imajo, Kentaro, et al.
Published: (2025)
by: Imajo, Kentaro, et al.
Published: (2025)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
by: Eigler, Lukáš, et al.
Published: (2026)
by: Eigler, Lukáš, et al.
Published: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
by: Yuan, Tongxin, et al.
Published: (2024)
by: Yuan, Tongxin, et al.
Published: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
by: Wu, Junjie, et al.
Published: (2026)
by: Wu, Junjie, et al.
Published: (2026)
Quantitative LLM Judges
by: Sahoo, Aishwarya, et al.
Published: (2025)
by: Sahoo, Aishwarya, et al.
Published: (2025)
How Reliable is Multilingual LLM-as-a-Judge?
by: Fu, Xiyan, et al.
Published: (2025)
by: Fu, Xiyan, et al.
Published: (2025)
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
by: Sternlicht, Noy, et al.
Published: (2025)
by: Sternlicht, Noy, et al.
Published: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
MR. Judge: Multimodal Reasoner as a Judge
by: Pi, Renjie, et al.
Published: (2025)
by: Pi, Renjie, et al.
Published: (2025)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
by: Le, Hieu Xuan, et al.
Published: (2026)
by: Le, Hieu Xuan, et al.
Published: (2026)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
by: Lee, Dongryeol, et al.
Published: (2026)
by: Lee, Dongryeol, et al.
Published: (2026)
Think-J: Learning to Think for Generative LLM-as-a-Judge
by: Huang, Hui, et al.
Published: (2025)
by: Huang, Hui, et al.
Published: (2025)
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
by: Yang, Jingbo, et al.
Published: (2026)
by: Yang, Jingbo, et al.
Published: (2026)
AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
by: Miao, Yongliang, et al.
Published: (2026)
by: Miao, Yongliang, et al.
Published: (2026)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
by: Thakur, Aman Singh, et al.
Published: (2024)
by: Thakur, Aman Singh, et al.
Published: (2024)
Similar Items
-
DRESS: Disentangled Representation-based Self-Supervised Meta-Learning for Diverse Tasks
by: Cui, Wei, et al.
Published: (2025) -
MSc-SQL: Multi-Sample Critiquing Small Language Models For Text-To-SQL Translation
by: Gorti, Satya Krishna, et al.
Published: (2024) -
Tabular Data Contrastive Learning via Class-Conditioned and Feature-Correlation Based Augmentation
by: Cui, Wei, et al.
Published: (2024) -
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
by: Tang, Yuqi, et al.
Published: (2025) -
A Geometric Framework for Understanding Memorization in Generative Models
by: Ross, Brendan Leigh, et al.
Published: (2024)