Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Yehudai, Asaf, Eden, Lilach, Shmueli-Scheuer, Michal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
by: Keren, Tomer, et al.
Published: (2026)
by: Keren, Tomer, et al.
Published: (2026)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025)
by: Kour, George, et al.
Published: (2025)
General Agent Evaluation
by: Bandel, Elron, et al.
Published: (2026)
by: Bandel, Elron, et al.
Published: (2026)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
by: Wang, Leyao, et al.
Published: (2026)
by: Wang, Leyao, et al.
Published: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
by: Habba, Eliya, et al.
Published: (2026)
by: Habba, Eliya, et al.
Published: (2026)
Robustness as an Emergent Property of Task Performance
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
WildIFEval: Instruction Following in the Wild
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
by: Halfon, Alon, et al.
Published: (2024)
by: Halfon, Alon, et al.
Published: (2024)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
by: Bandel, Elron, et al.
Published: (2024)
by: Bandel, Elron, et al.
Published: (2024)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models
by: Huber, Thomas, et al.
Published: (2025)
by: Huber, Thomas, et al.
Published: (2025)
CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation
by: Jiang, Yuyang, et al.
Published: (2025)
by: Jiang, Yuyang, et al.
Published: (2025)
AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
by: Ruan, Jianhao, et al.
Published: (2026)
by: Ruan, Jianhao, et al.
Published: (2026)
Zodiac: A Cardiologist-Level LLM Framework for Multi-Agent Diagnostics
by: Zhou, Yuan, et al.
Published: (2024)
by: Zhou, Yuan, et al.
Published: (2024)
CAMO: An Agentic Framework for Automated Causal Discovery from Micro Behaviors to Macro Emergence in LLM Agent Simulations
by: Yu, Xiangning, et al.
Published: (2026)
by: Yu, Xiangning, et al.
Published: (2026)
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
by: Gao, Joshua, et al.
Published: (2025)
by: Gao, Joshua, et al.
Published: (2025)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
by: Roy, Joyjit, et al.
Published: (2026)
by: Roy, Joyjit, et al.
Published: (2026)
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models
by: Gupta, Sonam, et al.
Published: (2024)
by: Gupta, Sonam, et al.
Published: (2024)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
by: Kartik, NVJK, et al.
Published: (2025)
by: Kartik, NVJK, et al.
Published: (2025)
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
by: Hu, Yuanzhe, et al.
Published: (2025)
by: Hu, Yuanzhe, et al.
Published: (2025)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
Automated Survey Collection with LLM-based Conversational Agents
by: Kaiyrbekov, Kurmanbek, et al.
Published: (2025)
by: Kaiyrbekov, Kurmanbek, et al.
Published: (2025)
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
by: Deng, Jiaqi, et al.
Published: (2025)
by: Deng, Jiaqi, et al.
Published: (2025)
AutoAgent: A Fully-Automated and Zero-Code Framework for LLM Agents
by: Tang, Jiabin, et al.
Published: (2025)
by: Tang, Jiabin, et al.
Published: (2025)
Efficient Benchmarking of Language Models
by: Perlitz, Yotam, et al.
Published: (2023)
by: Perlitz, Yotam, et al.
Published: (2023)
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
by: Li, Weizhen, et al.
Published: (2025)
by: Li, Weizhen, et al.
Published: (2025)
RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents
by: Rosati, Riccardo, et al.
Published: (2026)
by: Rosati, Riccardo, et al.
Published: (2026)
Toward Subtrait-Level Model Explainability in Automated Writing Evaluation
by: Andrade-Lotero, Alejandro, et al.
Published: (2025)
by: Andrade-Lotero, Alejandro, et al.
Published: (2025)
Competition-Level Problems are Effective LLM Evaluators
by: Huang, Yiming, et al.
Published: (2023)
by: Huang, Yiming, et al.
Published: (2023)
Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting
by: Meincke, Lennart, et al.
Published: (2025)
by: Meincke, Lennart, et al.
Published: (2025)
Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
by: Meincke, Lennart, et al.
Published: (2025)
by: Meincke, Lennart, et al.
Published: (2025)
Prompting Science Report 1: Prompt Engineering is Complicated and Contingent
by: Meincke, Lennart, et al.
Published: (2025)
by: Meincke, Lennart, et al.
Published: (2025)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
by: Komoravolu, Sameer, et al.
Published: (2025)
by: Komoravolu, Sameer, et al.
Published: (2025)
Similar Items
-
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025) -
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025) -
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024) -
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
by: Keren, Tomer, et al.
Published: (2026) -
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025)