Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Chong, Penny, Abichandani, Harshavardhan, Shen, Jiyuan, Ghosh, Atin, Moe, Min Pyae, Mai, Yifan, Dahlmeier, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
by: Shen, Jiyuan, et al.
Published: (2026)
by: Shen, Jiyuan, et al.
Published: (2026)
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
by: Sadjoli, Nicholas, et al.
Published: (2026)
by: Sadjoli, Nicholas, et al.
Published: (2026)
From Prompts to Worlds: How Users Iterate, Explore, and Make Sense of AI-Generated 3D Environments
by: Pyae, Aung
Published: (2026)
by: Pyae, Aung
Published: (2026)
Self-Anchoring Calibration Drift in Large Language Models: How Multi-Turn Conversations Reshape Model Confidence
by: Harshavardhan
Published: (2026)
by: Harshavardhan
Published: (2026)
Understanding Student Acceptance, Trust, and Attitudes Toward AI-Generated Images for Educational Purposes
by: Pyae, Aung
Published: (2024)
by: Pyae, Aung
Published: (2024)
Decision-aware User Simulation Agent for Evaluating Conversational Recommender Systems
by: Li, Yuan-Chi, et al.
Published: (2026)
by: Li, Yuan-Chi, et al.
Published: (2026)
HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
by: Jin, Chao, et al.
Published: (2026)
by: Jin, Chao, et al.
Published: (2026)
Nonstandard Errors in AI Agents
by: Gao, Ruijiang, et al.
Published: (2026)
by: Gao, Ruijiang, et al.
Published: (2026)
Evaluating Cognitive Age Alignment in Interactive AI Agents
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
Addressing Situated Teaching Needs: A Multi-Agent Framework for Automated Slide Adaptation
by: Liu, Binglin, et al.
Published: (2025)
by: Liu, Binglin, et al.
Published: (2025)
ARTEMIS-DA: An Advanced Reasoning and Transformation Engine for Multi-Step Insight Synthesis in Data Analytics
by: Hussain, Atin Sakkeer
Published: (2024)
by: Hussain, Atin Sakkeer
Published: (2024)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
by: Pothiraj, Atin, et al.
Published: (2025)
by: Pothiraj, Atin, et al.
Published: (2025)
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
by: Wu, Xinyi, et al.
Published: (2026)
by: Wu, Xinyi, et al.
Published: (2026)
Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents
by: Cheng, Liang, et al.
Published: (2026)
by: Cheng, Liang, et al.
Published: (2026)
Efficient Agent Evaluation via Diversity-Guided User Simulation
by: Nakash, Itay, et al.
Published: (2026)
by: Nakash, Itay, et al.
Published: (2026)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
by: Sirdeshmukh, Ved, et al.
Published: (2026)
by: Sirdeshmukh, Ved, et al.
Published: (2026)
Exploring Human-in-the-Loop Themes in AI Application Development: An Empirical Thematic Analysis
by: Suksakul, Parm, et al.
Published: (2026)
by: Suksakul, Parm, et al.
Published: (2026)
Interactively Diagnosing Errors in a Semantic Parser
by: Nakos, Constantine, et al.
Published: (2024)
by: Nakos, Constantine, et al.
Published: (2024)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation
by: Sawarni, Ayush, et al.
Published: (2026)
by: Sawarni, Ayush, et al.
Published: (2026)
Reliable Photoswitching of Vesicle‐Like Structures Assembled from Twisted Azo Derivatives
by: Pyae Thu, et al.
Published: (2025)
by: Pyae Thu, et al.
Published: (2025)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026)
by: Puyin, Li, et al.
Published: (2026)
Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
by: An, Subin, et al.
Published: (2025)
by: An, Subin, et al.
Published: (2025)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
by: Komoravolu, Sameer, et al.
Published: (2025)
by: Komoravolu, Sameer, et al.
Published: (2025)
Document Intelligence in the Era of Large Language Models: A Survey
by: Wang, Weishi, et al.
Published: (2025)
by: Wang, Weishi, et al.
Published: (2025)
StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error
by: Yang, Shu-Xun, et al.
Published: (2025)
by: Yang, Shu-Xun, et al.
Published: (2025)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents
by: Liang, Yijuan, et al.
Published: (2026)
by: Liang, Yijuan, et al.
Published: (2026)
Just aware enough: Evaluating awareness across artificial systems
by: Meertens, Nadine, et al.
Published: (2026)
by: Meertens, Nadine, et al.
Published: (2026)
$\texttt{PatentAgent}$: Intelligent Agent for Automated Pharmaceutical Patent Analysis
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Large Language Models for Judicial Entity Extraction: A Comparative Study
by: Hussain, Atin Sakkeer, et al.
Published: (2024)
by: Hussain, Atin Sakkeer, et al.
Published: (2024)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
by: Kapoor, Sayash, et al.
Published: (2025)
by: Kapoor, Sayash, et al.
Published: (2025)
Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation?
by: Li, Wensu, et al.
Published: (2026)
by: Li, Wensu, et al.
Published: (2026)
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
by: Ossowski, Timothy, et al.
Published: (2026)
by: Ossowski, Timothy, et al.
Published: (2026)
TalkToAgent: A Human-centric Explanation of Reinforcement Learning Agents with Large Language Models
by: Kim, Haechang, et al.
Published: (2025)
by: Kim, Haechang, et al.
Published: (2025)
Hierarchical Scoring for Machine Learning Classifier Error Impact Evaluation
by: Lanus, Erin, et al.
Published: (2025)
by: Lanus, Erin, et al.
Published: (2025)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
by: Kazi, Taaha, et al.
Published: (2024)
by: Kazi, Taaha, et al.
Published: (2024)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
Similar Items
-
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
by: Shen, Jiyuan, et al.
Published: (2026) -
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
by: Sadjoli, Nicholas, et al.
Published: (2026) -
From Prompts to Worlds: How Users Iterate, Explore, and Make Sense of AI-Generated 3D Environments
by: Pyae, Aung
Published: (2026) -
Self-Anchoring Calibration Drift in Large Language Models: How Multi-Turn Conversations Reshape Model Confidence
by: Harshavardhan
Published: (2026) -
Understanding Student Acceptance, Trust, and Attitudes Toward AI-Generated Images for Educational Purposes
by: Pyae, Aung
Published: (2024)