JAF: Judge Agent Forest
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garg, Sahil, Cheezum, Brad, Dutta, Sridhar, Agarwal, Vishal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Road of Adaptive AI for Precision in Cybersecurity
von: Garg, Sahil
Veröffentlicht: (2025)
von: Garg, Sahil
Veröffentlicht: (2025)
Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistants
von: Herrera, Alejandro Breen, et al.
Veröffentlicht: (2026)
von: Herrera, Alejandro Breen, et al.
Veröffentlicht: (2026)
What is in a name? Mitigating Name Bias in Text Embeddings via Anonymization
von: Manchanda, Sahil, et al.
Veröffentlicht: (2025)
von: Manchanda, Sahil, et al.
Veröffentlicht: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
von: Qu, Yuxiao, et al.
Veröffentlicht: (2024)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2024)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
ELSA: A Style Aligned Dataset for Emotionally Intelligent Language Generation
von: Gandhi, Vishal, et al.
Veröffentlicht: (2025)
von: Gandhi, Vishal, et al.
Veröffentlicht: (2025)
Local Prompt Optimization
von: Jain, Yash, et al.
Veröffentlicht: (2025)
von: Jain, Yash, et al.
Veröffentlicht: (2025)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
von: Jain, Dhruv, et al.
Veröffentlicht: (2025)
von: Jain, Dhruv, et al.
Veröffentlicht: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
The Perfect Blend: Redefining RLHF with Mixture of Judges
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
Investigating Non-Transitivity in LLM-as-a-Judge
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
CLewR: Curriculum Learning with Restarts for Machine Translation Preference Learning
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)
von: Dragomir, Alexandra, et al.
Veröffentlicht: (2026)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
Debate Helps Weak Judges Reward Stronger Models
von: Elasky, Ethan, et al.
Veröffentlicht: (2026)
von: Elasky, Ethan, et al.
Veröffentlicht: (2026)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
von: Singh, Aasheesh, et al.
Veröffentlicht: (2025)
von: Singh, Aasheesh, et al.
Veröffentlicht: (2025)
WebArena: A Realistic Web Environment for Building Autonomous Agents
von: Zhou, Shuyan, et al.
Veröffentlicht: (2023)
von: Zhou, Shuyan, et al.
Veröffentlicht: (2023)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
Context Over Content: Exposing Evaluation Faking in Automated Judges
von: Gupta, Manan, et al.
Veröffentlicht: (2026)
von: Gupta, Manan, et al.
Veröffentlicht: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
von: Gera, Ariel, et al.
Veröffentlicht: (2024)
von: Gera, Ariel, et al.
Veröffentlicht: (2024)
Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
von: Jwa, Seungyeon, et al.
Veröffentlicht: (2025)
von: Jwa, Seungyeon, et al.
Veröffentlicht: (2025)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
von: Salinas, David, et al.
Veröffentlicht: (2025)
von: Salinas, David, et al.
Veröffentlicht: (2025)
Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
von: Dissanayake, Pasan, et al.
Veröffentlicht: (2025)
von: Dissanayake, Pasan, et al.
Veröffentlicht: (2025)
Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
von: Agarwal, Rajan, et al.
Veröffentlicht: (2025)
von: Agarwal, Rajan, et al.
Veröffentlicht: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
von: Xu, Ruiyao, et al.
Veröffentlicht: (2026)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
von: Gupta, Manan, et al.
Veröffentlicht: (2026)
von: Gupta, Manan, et al.
Veröffentlicht: (2026)
Approximating Human Preferences Using a Multi-Judge Learned System
von: Sprejer, Eitán, et al.
Veröffentlicht: (2025)
von: Sprejer, Eitán, et al.
Veröffentlicht: (2025)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
von: Hu, Renjun, et al.
Veröffentlicht: (2025)
von: Hu, Renjun, et al.
Veröffentlicht: (2025)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
von: Liu, Zijun, et al.
Veröffentlicht: (2024)
von: Liu, Zijun, et al.
Veröffentlicht: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
von: Landesberg, Eddie
Veröffentlicht: (2026)
von: Landesberg, Eddie
Veröffentlicht: (2026)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
von: Zhang, Wenbo, et al.
Veröffentlicht: (2026)
von: Zhang, Wenbo, et al.
Veröffentlicht: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Road of Adaptive AI for Precision in Cybersecurity
von: Garg, Sahil
Veröffentlicht: (2025) -
Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistants
von: Herrera, Alejandro Breen, et al.
Veröffentlicht: (2026) -
What is in a name? Mitigating Name Bias in Text Embeddings via Anonymization
von: Manchanda, Sahil, et al.
Veröffentlicht: (2025) -
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026) -
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)