Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | AlShikh, Waseem, Ali, Muayad Sayed, Kennedy, Brian, Mozolevskyi, Dmytro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comparative Analysis of Retrieval Systems in the Real World
by: Mozolevskyi, Dmytro, et al.
Published: (2024)
by: Mozolevskyi, Dmytro, et al.
Published: (2024)
Expect the Unexpected: FailSafe Long Context QA for Finance
by: Kamble, Kiran, et al.
Published: (2025)
by: Kamble, Kiran, et al.
Published: (2025)
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
by: Bensal, Shelly, et al.
Published: (2025)
by: Bensal, Shelly, et al.
Published: (2025)
Understanding AI Evaluation Patterns: How Different GPT Models Assess Vision-Language Descriptions
by: Abdoli, Sajjad, et al.
Published: (2025)
by: Abdoli, Sajjad, et al.
Published: (2025)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
by: Mirtaheri, Mehrnoosh, et al.
Published: (2024)
by: Mirtaheri, Mehrnoosh, et al.
Published: (2024)
Towards More Standardized AI Evaluation: From Models to Agents
by: Filali, Ali El, et al.
Published: (2026)
by: Filali, Ali El, et al.
Published: (2026)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
by: Kazi, Taaha, et al.
Published: (2024)
by: Kazi, Taaha, et al.
Published: (2024)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
FLUKE: A Linguistically-Driven and Task-Agnostic Framework for Robustness Evaluation
by: Otmakhova, Yulia, et al.
Published: (2025)
by: Otmakhova, Yulia, et al.
Published: (2025)
Writing in the Margins: Better Inference Pattern for Long Context Retrieval
by: Russak, Melisa, et al.
Published: (2024)
by: Russak, Melisa, et al.
Published: (2024)
DARD: A Multi-Agent Approach for Task-Oriented Dialog Systems
by: Gupta, Aman, et al.
Published: (2024)
by: Gupta, Aman, et al.
Published: (2024)
Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
by: Shen, Yiting, et al.
Published: (2026)
by: Shen, Yiting, et al.
Published: (2026)
Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
by: Yoon, Yejin, et al.
Published: (2025)
by: Yoon, Yejin, et al.
Published: (2025)
Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk
by: Ulmer, Dennis, et al.
Published: (2024)
by: Ulmer, Dennis, et al.
Published: (2024)
Controllable and Reliable Knowledge-Intensive Task-Oriented Conversational Agents with Declarative Genie Worksheets
by: Joshi, Harshit, et al.
Published: (2024)
by: Joshi, Harshit, et al.
Published: (2024)
Towards Goal-Oriented Agents for Evolving Problems Observed via Conversation
by: Free, Michael, et al.
Published: (2024)
by: Free, Michael, et al.
Published: (2024)
STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
by: Chhetri, Tek Raj, et al.
Published: (2025)
by: Chhetri, Tek Raj, et al.
Published: (2025)
Emergent Introspection in AI is Content-Agnostic
by: Lederman, Harvey, et al.
Published: (2026)
by: Lederman, Harvey, et al.
Published: (2026)
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
by: Yang, Ke, et al.
Published: (2026)
by: Yang, Ke, et al.
Published: (2026)
InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
by: Hu, Xueyu, et al.
Published: (2024)
by: Hu, Xueyu, et al.
Published: (2024)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
by: Liu, Zijun, et al.
Published: (2023)
by: Liu, Zijun, et al.
Published: (2023)
Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
by: Shu, Raphael, et al.
Published: (2024)
by: Shu, Raphael, et al.
Published: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
by: Baidya, Avinash, et al.
Published: (2025)
by: Baidya, Avinash, et al.
Published: (2025)
SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
by: Si, Shuzheng, et al.
Published: (2023)
by: Si, Shuzheng, et al.
Published: (2023)
SPRING Lab IITM's submission to Low Resource Indic Language Translation Shared Task
by: Sayed, Hamees, et al.
Published: (2024)
by: Sayed, Hamees, et al.
Published: (2024)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
by: Kapoor, Sayash, et al.
Published: (2025)
by: Kapoor, Sayash, et al.
Published: (2025)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
by: Kinniment, Megan, et al.
Published: (2023)
by: Kinniment, Megan, et al.
Published: (2023)
Can Large Language Models Predict the Outcome of Judicial Decisions?
by: Kmainasi, Mohamed Bayan, et al.
Published: (2025)
by: Kmainasi, Mohamed Bayan, et al.
Published: (2025)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
by: Zhang, Yifei, et al.
Published: (2026)
by: Zhang, Yifei, et al.
Published: (2026)
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
FamiCom: Further Demystifying Prompts for Language Models with Task-Agnostic Performance Estimation
by: Li, Bangzheng, et al.
Published: (2024)
by: Li, Bangzheng, et al.
Published: (2024)
Holistic Evaluation and Failure Diagnosis of AI Agents
by: Madvil, Netta, et al.
Published: (2026)
by: Madvil, Netta, et al.
Published: (2026)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
by: Komoravolu, Sameer, et al.
Published: (2025)
by: Komoravolu, Sameer, et al.
Published: (2025)
Mind the Goal: Data-Efficient Goal-Oriented Evaluation of Conversational Agents and Chatbots using Teacher Models
by: Piskala, Deepak Babu, et al.
Published: (2025)
by: Piskala, Deepak Babu, et al.
Published: (2025)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
by: Emery, Deanna, et al.
Published: (2025)
by: Emery, Deanna, et al.
Published: (2025)
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
by: Li, Minghao, et al.
Published: (2025)
by: Li, Minghao, et al.
Published: (2025)
First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
by: Vitel, Dmytro, et al.
Published: (2025)
by: Vitel, Dmytro, et al.
Published: (2025)
Similar Items
-
Comparative Analysis of Retrieval Systems in the Real World
by: Mozolevskyi, Dmytro, et al.
Published: (2024) -
Expect the Unexpected: FailSafe Long Context QA for Finance
by: Kamble, Kiran, et al.
Published: (2025) -
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
by: Bensal, Shelly, et al.
Published: (2025) -
Understanding AI Evaluation Patterns: How Different GPT Models Assess Vision-Language Descriptions
by: Abdoli, Sajjad, et al.
Published: (2025) -
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
by: Mirtaheri, Mehrnoosh, et al.
Published: (2024)