Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Jiaju, Lu, Yuxuan, Wang, Xiaojie, Zeng, Huimin, Huang, Jing, Gesi, Jiri, Xu, Ying, Yao, Bingsheng, Wang, Dakuo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LangMARL: Natural Language Multi-Agent Reinforcement Learning
by: Yao, Huaiyuan, et al.
Published: (2026)
by: Yao, Huaiyuan, et al.
Published: (2026)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
Pitfalls in Evaluating Interpretability Agents
by: Haklay, Tal, et al.
Published: (2026)
by: Haklay, Tal, et al.
Published: (2026)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
by: Xi, Wang, et al.
Published: (2025)
by: Xi, Wang, et al.
Published: (2025)
ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions
by: Gupta, Aayush
Published: (2026)
by: Gupta, Aayush
Published: (2026)
BMAM: Brain-inspired Multi-Agent Memory Framework
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
by: Fan, Jingxing, et al.
Published: (2025)
by: Fan, Jingxing, et al.
Published: (2025)
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
by: Zong, Chang, et al.
Published: (2024)
by: Zong, Chang, et al.
Published: (2024)
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
by: Cheng, Ziming, et al.
Published: (2025)
by: Cheng, Ziming, et al.
Published: (2025)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
by: Xu, Weijie, et al.
Published: (2024)
by: Xu, Weijie, et al.
Published: (2024)
MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional
by: Cruz, Roberto, et al.
Published: (2026)
by: Cruz, Roberto, et al.
Published: (2026)
ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary
by: Li, Yutong, et al.
Published: (2024)
by: Li, Yutong, et al.
Published: (2024)
AI Agents-as-Judge: Automated Assessment of Accuracy, Consistency, Completeness and Clarity for Enterprise Documents
by: Dasgupta, Sudip, et al.
Published: (2025)
by: Dasgupta, Sudip, et al.
Published: (2025)
AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement
by: Rosser, J, et al.
Published: (2025)
by: Rosser, J, et al.
Published: (2025)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
by: Anam, Rizal Khoirul
Published: (2025)
by: Anam, Rizal Khoirul
Published: (2025)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
by: Garg, Saloni, et al.
Published: (2026)
by: Garg, Saloni, et al.
Published: (2026)
MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
by: Lim, Ernest, et al.
Published: (2025)
by: Lim, Ernest, et al.
Published: (2025)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
by: Jia, Xiao
Published: (2026)
by: Jia, Xiao
Published: (2026)
XPath Agent: An Efficient XPath Programming Agent Based on LLM for Web Crawler
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
by: Aksoy, Sinan G., et al.
Published: (2026)
by: Aksoy, Sinan G., et al.
Published: (2026)
LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering
by: Liu, Beiming, et al.
Published: (2025)
by: Liu, Beiming, et al.
Published: (2025)
Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
by: Otal, Hakan T., et al.
Published: (2024)
by: Otal, Hakan T., et al.
Published: (2024)
RecAI: Leveraging Large Language Models for Next-Generation Recommender Systems
by: Lian, Jianxun, et al.
Published: (2024)
by: Lian, Jianxun, et al.
Published: (2024)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
An NLP-Driven Framework for Curriculum-Labor Market Alignment: Schema-Constrained LLM Extraction, ESCO-Anchored Semantic Matching, and Multi-Dimensional Gap Quantification
by: Turaev, Sherzod, et al.
Published: (2026)
by: Turaev, Sherzod, et al.
Published: (2026)
Aligning Large Language Models for Controllable Recommendations
by: Lu, Wensheng, et al.
Published: (2024)
by: Lu, Wensheng, et al.
Published: (2024)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
by: Yao, Ben, et al.
Published: (2025)
by: Yao, Ben, et al.
Published: (2025)
Optimizing What We Trust: Reliability-Guided QUBO Selection of Multi-Agent Weak Framing Signals for Arabic Sentiment Prediction
by: Alkhalifa, Rabab
Published: (2026)
by: Alkhalifa, Rabab
Published: (2026)
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
SepsisLab: Early Sepsis Prediction with Uncertainty Quantification and Active Sensing
by: Yin, Changchang, et al.
Published: (2024)
by: Yin, Changchang, et al.
Published: (2024)
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
by: Yu, Zhiqi, et al.
Published: (2026)
by: Yu, Zhiqi, et al.
Published: (2026)
LLMs as Deceptive Agents: How Role-Based Prompting Induces Semantic Ambiguity in Puzzle Tasks
by: Yoo, Seunghyun
Published: (2025)
by: Yoo, Seunghyun
Published: (2025)
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
by: Wang, Yongjie, et al.
Published: (2025)
by: Wang, Yongjie, et al.
Published: (2025)
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
by: Ferdousi, Rahatara, et al.
Published: (2025)
by: Ferdousi, Rahatara, et al.
Published: (2025)
A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query Processing
by: Nourmohammadi, Naeimeh, et al.
Published: (2026)
by: Nourmohammadi, Naeimeh, et al.
Published: (2026)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
by: Ikoma, Hayato, et al.
Published: (2025)
by: Ikoma, Hayato, et al.
Published: (2025)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
by: Khan, Omer Jauhar
Published: (2025)
by: Khan, Omer Jauhar
Published: (2025)
Similar Items
-
LangMARL: Natural Language Multi-Agent Reinforcement Learning
by: Yao, Huaiyuan, et al.
Published: (2026) -
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025) -
Pitfalls in Evaluating Interpretability Agents
by: Haklay, Tal, et al.
Published: (2026) -
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
by: Lu, Yuxuan, et al.
Published: (2025) -
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
by: Xi, Wang, et al.
Published: (2025)