ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yifei, Nayyeri, Hooshang, Khaziev, Rinat, Yilmaz, Emine, Tur, Gokhan, Hakkani-Tür, Dilek, Thadakamalla, Hari |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
by: Dongre, Vardhan, et al.
Published: (2026)
by: Dongre, Vardhan, et al.
Published: (2026)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
Instruct, Not Assist: LLM-based Multi-Turn Planning and Hierarchical Questioning for Socratic Code Debugging
by: Kargupta, Priyanka, et al.
Published: (2024)
by: Kargupta, Priyanka, et al.
Published: (2024)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
by: Kazi, Taaha, et al.
Published: (2024)
by: Kazi, Taaha, et al.
Published: (2024)
MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
by: Dongre, Vardhan, et al.
Published: (2025)
by: Dongre, Vardhan, et al.
Published: (2025)
Know Your Mistakes: Towards Preventing Overreliance on Task-Oriented Conversational AI Through Accountability Modeling
by: Dey, Suvodip, et al.
Published: (2025)
by: Dey, Suvodip, et al.
Published: (2025)
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
by: Kasprova, Vira, et al.
Published: (2026)
by: Kasprova, Vira, et al.
Published: (2026)
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Confidence Estimation for LLM-Based Dialogue State Tracking
by: Sun, Yi-Jyun, et al.
Published: (2024)
by: Sun, Yi-Jyun, et al.
Published: (2024)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
by: Rabbani, Parisa, et al.
Published: (2025)
by: Rabbani, Parisa, et al.
Published: (2025)
Towards Self-Improving Error Diagnosis in Multi-Agent Systems
by: Li, Jiazheng, et al.
Published: (2026)
by: Li, Jiazheng, et al.
Published: (2026)
Simulating User Agents for Embodied Conversational-AI
by: Philipov, Daniel, et al.
Published: (2024)
by: Philipov, Daniel, et al.
Published: (2024)
Adaptive Multi-Agent Response Refinement in Conversational Systems
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
by: Bogavelli, Tara, et al.
Published: (2025)
by: Bogavelli, Tara, et al.
Published: (2025)
YourBench: Easy Custom Evaluation Sets for Everyone
by: Shashidhar, Sumuk, et al.
Published: (2025)
by: Shashidhar, Sumuk, et al.
Published: (2025)
Plan Verification for LLM-Based Embodied Task Completion Agents
by: Hariharan, Ananth, et al.
Published: (2025)
by: Hariharan, Ananth, et al.
Published: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
by: Singh, Janvijay, et al.
Published: (2026)
by: Singh, Janvijay, et al.
Published: (2026)
Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning Data
by: Agarwal, Ishika, et al.
Published: (2025)
by: Agarwal, Ishika, et al.
Published: (2025)
Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation
by: Harrasse, Abir, et al.
Published: (2024)
by: Harrasse, Abir, et al.
Published: (2024)
Self-Improving LLM Agents at Test-Time
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Empowering LLMs in Task-Oriented Dialogues: A Domain-Independent Multi-Agent Framework and Fine-Tuning Strategy
by: Feng, Zihao, et al.
Published: (2025)
by: Feng, Zihao, et al.
Published: (2025)
Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems
by: İnan, Mert, et al.
Published: (2025)
by: İnan, Mert, et al.
Published: (2025)
AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
Goal Alignment in LLM-Based User Simulators for Conversational AI
by: Mehri, Shuhaib, et al.
Published: (2025)
by: Mehri, Shuhaib, et al.
Published: (2025)
Adaptive Monitoring and Real-World Evaluation of Agentic AI Systems
by: Shukla, Manish
Published: (2025)
by: Shukla, Manish
Published: (2025)
Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integration and Evaluation using Novel Metrics and Dataset
by: Gabriel, Adrian Garret, et al.
Published: (2024)
by: Gabriel, Adrian Garret, et al.
Published: (2024)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
GNNs as Predictors of Agentic Workflow Performances
by: Zhang, Yuanshuo, et al.
Published: (2025)
by: Zhang, Yuanshuo, et al.
Published: (2025)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
by: Liu, Zijun, et al.
Published: (2023)
by: Liu, Zijun, et al.
Published: (2023)
System of Agentic AI for the Discovery of Metal-Organic Frameworks
by: Inizan, Theo Jaffrelot, et al.
Published: (2025)
by: Inizan, Theo Jaffrelot, et al.
Published: (2025)
SPASM: Stable Persona-driven Agent Simulation for Multi-turn Dialogue Generation
by: Luo, Han, et al.
Published: (2026)
by: Luo, Han, et al.
Published: (2026)
Exploring Modularity of Agentic Systems for Drug Discovery
by: van Weesep, Laura, et al.
Published: (2025)
by: van Weesep, Laura, et al.
Published: (2025)
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
by: Li, Shiyang, et al.
Published: (2026)
by: Li, Shiyang, et al.
Published: (2026)
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
by: Akshathala, Sreemaee, et al.
Published: (2025)
by: Akshathala, Sreemaee, et al.
Published: (2025)
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
ReSpAct: Harmonizing Reasoning, Speaking, and Acting Towards Building Large Language Model-Based Conversational AI Agents
by: Dongre, Vardhan, et al.
Published: (2024)
by: Dongre, Vardhan, et al.
Published: (2024)
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
by: Xu, Shihao, et al.
Published: (2026)
by: Xu, Shihao, et al.
Published: (2026)
Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
by: Zhang, Shaokun, et al.
Published: (2025)
by: Zhang, Shaokun, et al.
Published: (2025)
Hallucination Mitigation using Agentic AI Natural Language-Based Frameworks
by: Gosmar, Diego, et al.
Published: (2025)
by: Gosmar, Diego, et al.
Published: (2025)
A-MapReduce: Executing Wide Search via Agentic MapReduce
by: Chen, Mingju, et al.
Published: (2026)
by: Chen, Mingju, et al.
Published: (2026)
Similar Items
-
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
by: Dongre, Vardhan, et al.
Published: (2026) -
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
by: Bozdag, Nimet Beyza, et al.
Published: (2025) -
Instruct, Not Assist: LLM-based Multi-Turn Planning and Hierarchical Questioning for Socratic Code Debugging
by: Kargupta, Priyanka, et al.
Published: (2024) -
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
by: Kazi, Taaha, et al.
Published: (2024) -
MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
by: Dongre, Vardhan, et al.
Published: (2025)