Towards Automatic Evaluation of Task-Oriented Dialogue Flows
Fuente:
arXiv
Saved in:
| Main Authors: | Mirtaheri, Mehrnoosh, Varghese, Nikhil, Khatri, Chandra, Kelkar, Amol |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KULCQ: An Unsupervised Keyword-based Utterance Level Clustering Quality Metric
by: Guruprasad, Pranav, et al.
Published: (2024)
by: Guruprasad, Pranav, et al.
Published: (2024)
DiagGPT: An LLM-based and Multi-agent Dialogue System with Automatic Topic Management for Flexible Task-Oriented Dialogue
by: Cao, Lang
Published: (2023)
by: Cao, Lang
Published: (2023)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
by: Kazi, Taaha, et al.
Published: (2024)
by: Kazi, Taaha, et al.
Published: (2024)
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
by: Mirtaheri, Parsa, et al.
Published: (2026)
by: Mirtaheri, Parsa, et al.
Published: (2026)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
by: Zhang, Yifei, et al.
Published: (2026)
by: Zhang, Yifei, et al.
Published: (2026)
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
by: AlShikh, Waseem, et al.
Published: (2025)
by: AlShikh, Waseem, et al.
Published: (2025)
Unsupervised Flow Discovery from Task-oriented Dialogues
by: Ferreira, Patrícia, et al.
Published: (2024)
by: Ferreira, Patrícia, et al.
Published: (2024)
HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals
by: Mo, Lingbo, et al.
Published: (2024)
by: Mo, Lingbo, et al.
Published: (2024)
Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk
by: Ulmer, Dennis, et al.
Published: (2024)
by: Ulmer, Dennis, et al.
Published: (2024)
Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
by: Yoon, Yejin, et al.
Published: (2025)
by: Yoon, Yejin, et al.
Published: (2025)
LLM-Driven Multi-Turn Task-Oriented Dialogue Synthesis for Realistic Reasoning
by: Zhu, Yu, et al.
Published: (2026)
by: Zhu, Yu, et al.
Published: (2026)
Universal Post-Processing Networks for Joint Optimization of Modules in Task-Oriented Dialogue Systems
by: Ohashi, Atsumoto, et al.
Published: (2025)
by: Ohashi, Atsumoto, et al.
Published: (2025)
Decoupling Strategy and Execution in Task-Focused Dialogue via Goal-Oriented Preference Optimization
by: Xu, Jingyi, et al.
Published: (2026)
by: Xu, Jingyi, et al.
Published: (2026)
Analysis of Utterance Embeddings and Clustering Methods Related to Intent Induction for Task-Oriented Dialogue
by: Park, Jeiyoon, et al.
Published: (2022)
by: Park, Jeiyoon, et al.
Published: (2022)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
by: Ryan, Michael J., et al.
Published: (2025)
by: Ryan, Michael J., et al.
Published: (2025)
Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests
by: Mannekote, Amogh, et al.
Published: (2024)
by: Mannekote, Amogh, et al.
Published: (2024)
JMultiWOZ: A Large-Scale Japanese Multi-Domain Task-Oriented Dialogue Dataset
by: Ohashi, Atsumoto, et al.
Published: (2024)
by: Ohashi, Atsumoto, et al.
Published: (2024)
DuetSim: Building User Simulator with Dual Large Language Models for Task-Oriented Dialogues
by: Luo, Xiang, et al.
Published: (2024)
by: Luo, Xiang, et al.
Published: (2024)
Leveraging Graph Structures and Large Language Models for End-to-End Synthetic Task-Oriented Dialogues
by: Medjad, Maya, et al.
Published: (2025)
by: Medjad, Maya, et al.
Published: (2025)
Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings
by: Elizabeth, Michelle, et al.
Published: (2024)
by: Elizabeth, Michelle, et al.
Published: (2024)
Many Hands Make Light Work: Task-Oriented Dialogue System with Module-Based Mixture-of-Experts
by: Su, Ruolin, et al.
Published: (2024)
by: Su, Ruolin, et al.
Published: (2024)
SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
by: Si, Shuzheng, et al.
Published: (2023)
by: Si, Shuzheng, et al.
Published: (2023)
Decision-Oriented Dialogue for Human-AI Collaboration
by: Lin, Jessy, et al.
Published: (2023)
by: Lin, Jessy, et al.
Published: (2023)
STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media
by: Xue, Liang, et al.
Published: (2026)
by: Xue, Liang, et al.
Published: (2026)
MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues
by: Xue, Liang, et al.
Published: (2025)
by: Xue, Liang, et al.
Published: (2025)
Towards Negotiative Dialogue for the Talkamatic Dialogue Manager
by: Larsson, Staffan, et al.
Published: (2024)
by: Larsson, Staffan, et al.
Published: (2024)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
Anti-Overestimation Dialogue Policy Learning for Task-Completion Dialogue System
by: Tian, Chang, et al.
Published: (2022)
by: Tian, Chang, et al.
Published: (2022)
Beyond Ontology in Dialogue State Tracking for Goal-Oriented Chatbot
by: Lee, Sejin, et al.
Published: (2024)
by: Lee, Sejin, et al.
Published: (2024)
Improving Multi-Domain Task-Oriented Dialogue System with Offline Reinforcement Learning
by: Prajapat, Dharmendra, et al.
Published: (2024)
by: Prajapat, Dharmendra, et al.
Published: (2024)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
by: Kartik, NVJK, et al.
Published: (2025)
by: Kartik, NVJK, et al.
Published: (2025)
MADial-Bench: Towards Real-world Evaluation of Memory-Augmented Dialogue Generation
by: He, Junqing, et al.
Published: (2024)
by: He, Junqing, et al.
Published: (2024)
"In Dialogues We Learn": Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue Learning
by: Cheng, Chuanqi, et al.
Published: (2024)
by: Cheng, Chuanqi, et al.
Published: (2024)
Longitudinal Abuse and Sentiment Analysis of Hollywood Movie Dialogues using Language Models
by: Chandra, Rohitash, et al.
Published: (2025)
by: Chandra, Rohitash, et al.
Published: (2025)
Efficient Generation of Parameterised Quantum Circuits from Large Texts
by: Krawchuk, Colin, et al.
Published: (2025)
by: Krawchuk, Colin, et al.
Published: (2025)
METRO: Towards Strategy Induction from Expert Dialogue Transcripts for Non-collaborative Dialogues
by: Yang, Haofu, et al.
Published: (2026)
by: Yang, Haofu, et al.
Published: (2026)
Data Augmentation Integrating Dialogue Flow and Style to Adapt Spoken Dialogue Systems to Low-Resource User Groups
by: Qi, Zhiyang, et al.
Published: (2024)
by: Qi, Zhiyang, et al.
Published: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
DFlow: Diverse Dialogue Flow Simulation with Large Language Models
by: Du, Wanyu, et al.
Published: (2024)
by: Du, Wanyu, et al.
Published: (2024)
Similar Items
-
KULCQ: An Unsupervised Keyword-based Utterance Level Clustering Quality Metric
by: Guruprasad, Pranav, et al.
Published: (2024) -
DiagGPT: An LLM-based and Multi-agent Dialogue System with Automatic Topic Management for Flexible Task-Oriented Dialogue
by: Cao, Lang
Published: (2023) -
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
by: Kazi, Taaha, et al.
Published: (2024) -
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
by: Acikgoz, Emre Can, et al.
Published: (2025) -
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
by: Mirtaheri, Parsa, et al.
Published: (2026)