Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chhabra, Mukul, Medrano, Luigi, Verma, Arush |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Retrieval Augmented Generation with RAG Fusion: Lessons from an Industry Deployment
von: Medrano, Luigi, et al.
Veröffentlicht: (2026)
von: Medrano, Luigi, et al.
Veröffentlicht: (2026)
A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test
von: Sartori, Camilo Chacón, et al.
Veröffentlicht: (2026)
von: Sartori, Camilo Chacón, et al.
Veröffentlicht: (2026)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
von: Tsujimura, Hikaru, et al.
Veröffentlicht: (2025)
von: Tsujimura, Hikaru, et al.
Veröffentlicht: (2025)
Quality-Aware Translation Tagging in Multilingual RAG system
von: Moon, Hoyeon, et al.
Veröffentlicht: (2025)
von: Moon, Hoyeon, et al.
Veröffentlicht: (2025)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
von: Muhamed, Aashiq
Veröffentlicht: (2025)
von: Muhamed, Aashiq
Veröffentlicht: (2025)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
von: Monjur, Ocean, et al.
Veröffentlicht: (2026)
von: Monjur, Ocean, et al.
Veröffentlicht: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
von: Enguehard, Joseph, et al.
Veröffentlicht: (2025)
von: Enguehard, Joseph, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
von: Verga, Pat, et al.
Veröffentlicht: (2024)
von: Verga, Pat, et al.
Veröffentlicht: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
Routine: A Structural Planning Framework for LLM Agent System in Enterprise
von: Zeng, Guancheng, et al.
Veröffentlicht: (2025)
von: Zeng, Guancheng, et al.
Veröffentlicht: (2025)
Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems
von: Avinash, Karthik, et al.
Veröffentlicht: (2025)
von: Avinash, Karthik, et al.
Veröffentlicht: (2025)
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
von: Lin, Wei-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Wei-Hsiang, et al.
Veröffentlicht: (2025)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
von: Li, Weiyue, et al.
Veröffentlicht: (2026)
von: Li, Weiyue, et al.
Veröffentlicht: (2026)
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling
von: Yang, Jeff, et al.
Veröffentlicht: (2025)
von: Yang, Jeff, et al.
Veröffentlicht: (2025)
Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs
von: Singh, Gundeep, et al.
Veröffentlicht: (2026)
von: Singh, Gundeep, et al.
Veröffentlicht: (2026)
Knowledge-Graph Based RAG System Evaluation Framework
von: Dong, Sicheng, et al.
Veröffentlicht: (2025)
von: Dong, Sicheng, et al.
Veröffentlicht: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
von: Jin, Jiho, et al.
Veröffentlicht: (2026)
von: Jin, Jiho, et al.
Veröffentlicht: (2026)
CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades
von: Chang, Raeyoung, et al.
Veröffentlicht: (2026)
von: Chang, Raeyoung, et al.
Veröffentlicht: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
von: Li, Yihan, et al.
Veröffentlicht: (2025)
von: Li, Yihan, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Time Series Explanations
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
von: Sun, Bian, et al.
Veröffentlicht: (2026)
von: Sun, Bian, et al.
Veröffentlicht: (2026)
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2026)
von: Gisserot-Boukhlef, Hippolyte, et al.
Veröffentlicht: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
Are We on the Right Way to Assessing LLM-as-a-Judge?
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scaling Retrieval Augmented Generation with RAG Fusion: Lessons from an Industry Deployment
von: Medrano, Luigi, et al.
Veröffentlicht: (2026) -
A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test
von: Sartori, Camilo Chacón, et al.
Veröffentlicht: (2026) -
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
von: Tsujimura, Hikaru, et al.
Veröffentlicht: (2025) -
Quality-Aware Translation Tagging in Multilingual RAG system
von: Moon, Hoyeon, et al.
Veröffentlicht: (2025) -
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
von: Muhamed, Aashiq
Veröffentlicht: (2025)