Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs
Fuente:
arXiv
Saved in:
| Main Authors: | Komma, Abishek, Chandrasekarasastry, Nagesh Panyam, Leffel, Timothy, Goyal, Anuj, Metallinou, Angeliki, Matsoukas, Spyros, Galstyan, Aram |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging LLMs for Dialogue Quality Measurement
by: Jia, Jinghan, et al.
Published: (2024)
by: Jia, Jinghan, et al.
Published: (2024)
Training Zero-Shot Generalizable End-to-End Task-Oriented Dialog System Without Turn-level Dialog Annotations
by: Mosharrof, Adib, et al.
Published: (2024)
by: Mosharrof, Adib, et al.
Published: (2024)
Evaluating and Enhancing Out-of-Domain Generalization of Task-Oriented Dialog Systems for Task Completion without Turn-level Dialog Annotations
by: Mosharrof, Adib, et al.
Published: (2025)
by: Mosharrof, Adib, et al.
Published: (2025)
Zero-Shot Generalizable End-to-End Task-Oriented Dialog System using Context Summarization and Domain Schema
by: Mosharrof, Adib, et al.
Published: (2023)
by: Mosharrof, Adib, et al.
Published: (2023)
TOAD: Task-Oriented Automatic Dialogs with Diverse Response Styles
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
by: Markowitz, Elan, et al.
Published: (2025)
by: Markowitz, Elan, et al.
Published: (2025)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Task-Oriented Dialog Systems for the Senegalese Wolof Language
by: Mbaye, Derguene, et al.
Published: (2024)
by: Mbaye, Derguene, et al.
Published: (2024)
Actionable Conversational Quality Indicators for Improving Task-Oriented Dialog Systems
by: Higgins, Michael, et al.
Published: (2021)
by: Higgins, Michael, et al.
Published: (2021)
DARD: A Multi-Agent Approach for Task-Oriented Dialog Systems
by: Gupta, Aman, et al.
Published: (2024)
by: Gupta, Aman, et al.
Published: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
by: Baidya, Avinash, et al.
Published: (2025)
by: Baidya, Avinash, et al.
Published: (2025)
Conversation Routines: A Prompt Engineering Framework for Task-Oriented Dialog Systems
by: Robino, Giorgio
Published: (2025)
by: Robino, Giorgio
Published: (2025)
Comparing Data Augmentation Methods for End-to-End Task-Oriented Dialog Systems
by: Vlachos, Christos, et al.
Published: (2024)
by: Vlachos, Christos, et al.
Published: (2024)
On the steerability of large language models toward data-driven personas
by: Li, Junyi, et al.
Published: (2023)
by: Li, Junyi, et al.
Published: (2023)
ClarQ-LLM: A Benchmark for Models Clarifying and Requesting Information in Task-Oriented Dialog
by: Gan, Yujian, et al.
Published: (2024)
by: Gan, Yujian, et al.
Published: (2024)
Improving Multi-turn Task Completion in Task-Oriented Dialog Systems via Prompt Chaining and Fine-Grained Feedback
by: Fereidouni, Moghis, et al.
Published: (2025)
by: Fereidouni, Moghis, et al.
Published: (2025)
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
by: Stephen, Abishek, et al.
Published: (2026)
by: Stephen, Abishek, et al.
Published: (2026)
TA&AT: Enhancing Task-Oriented Dialog with Turn-Level Auxiliary Tasks and Action-Tree Based Scheduled Sampling
by: Liu, Longxiang, et al.
Published: (2024)
by: Liu, Longxiang, et al.
Published: (2024)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
by: Mirtaheri, Mehrnoosh, et al.
Published: (2024)
by: Mirtaheri, Mehrnoosh, et al.
Published: (2024)
Faithful Model Evaluation for Model-Based Metrics
by: Goyal, Palash, et al.
Published: (2023)
by: Goyal, Palash, et al.
Published: (2023)
Attribute Controlled Fine-tuning for Large Language Models: A Case Study on Detoxification
by: Meng, Tao, et al.
Published: (2024)
by: Meng, Tao, et al.
Published: (2024)
"Stupid robot, I want to speak to a human!" User Frustration Detection in Task-Oriented Dialog Systems
by: Caralt, Mireia Hernandez, et al.
Published: (2024)
by: Caralt, Mireia Hernandez, et al.
Published: (2024)
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
by: AlShikh, Waseem, et al.
Published: (2025)
by: AlShikh, Waseem, et al.
Published: (2025)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Granular Change Accuracy: A More Accurate Performance Metric for Dialogue State Tracking
by: Aksu, Taha, et al.
Published: (2024)
by: Aksu, Taha, et al.
Published: (2024)
Long Dialog Summarization: An Analysis
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
FANTAstic SEquences and Where to Find Them: Faithful and Efficient API Call Generation through State-tracked Constrained Decoding and Reranking
by: Wang, Zhuoer, et al.
Published: (2024)
by: Wang, Zhuoer, et al.
Published: (2024)
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
by: Hao, Tianxiang, et al.
Published: (2023)
by: Hao, Tianxiang, et al.
Published: (2023)
ESAinsTOD: A Unified End-to-End Schema-Aware Instruction-Tuning Framework for Task-Oriented Dialog Modeling
by: Teng, Dechuan, et al.
Published: (2026)
by: Teng, Dechuan, et al.
Published: (2026)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023)
by: Ovalle, Anaelia, et al.
Published: (2023)
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
by: Kumarage, Tharindu, et al.
Published: (2025)
by: Kumarage, Tharindu, et al.
Published: (2025)
Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems
by: Mishra, Sandeep, et al.
Published: (2025)
by: Mishra, Sandeep, et al.
Published: (2025)
Multitask Fine-Tuning and Generative Adversarial Learning for Improved Auxiliary Classification
by: Sun, Christopher, et al.
Published: (2024)
by: Sun, Christopher, et al.
Published: (2024)
BioMistral-NLU: Towards More Generalizable Medical Language Understanding through Instruction Tuning
by: Fu, Yujuan Velvin, et al.
Published: (2024)
by: Fu, Yujuan Velvin, et al.
Published: (2024)
A Comparative Empirical Study of Catastrophic Forgetting Mitigation in Sequential Task Adaptation for Continual Natural Language Processing Systems
by: Abrahamyan, Aram, et al.
Published: (2026)
by: Abrahamyan, Aram, et al.
Published: (2026)
Multi-dimensional Evaluation of Empathetic Dialog Responses
by: Xu, Zhichao, et al.
Published: (2024)
by: Xu, Zhichao, et al.
Published: (2024)
Towards Zero-Shot, Controllable Dialog Planning with LLMs
by: Väth, Dirk, et al.
Published: (2024)
by: Väth, Dirk, et al.
Published: (2024)
A LLM Benchmark based on the Minecraft Builder Dialog Agent Task
by: Madge, Chris, et al.
Published: (2024)
by: Madge, Chris, et al.
Published: (2024)
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
by: Song, Woomin, et al.
Published: (2025)
by: Song, Woomin, et al.
Published: (2025)
Learning from Relevant Subgoals in Successful Dialogs using Iterative Training for Task-oriented Dialog Systems
by: Kaiser, Magdalena, et al.
Published: (2024)
by: Kaiser, Magdalena, et al.
Published: (2024)
Similar Items
-
Leveraging LLMs for Dialogue Quality Measurement
by: Jia, Jinghan, et al.
Published: (2024) -
Training Zero-Shot Generalizable End-to-End Task-Oriented Dialog System Without Turn-level Dialog Annotations
by: Mosharrof, Adib, et al.
Published: (2024) -
Evaluating and Enhancing Out-of-Domain Generalization of Task-Oriented Dialog Systems for Task Completion without Turn-level Dialog Annotations
by: Mosharrof, Adib, et al.
Published: (2025) -
Zero-Shot Generalizable End-to-End Task-Oriented Dialog System using Context Summarization and Domain Schema
by: Mosharrof, Adib, et al.
Published: (2023) -
TOAD: Task-Oriented Automatic Dialogs with Diverse Response Styles
by: Liu, Yinhong, et al.
Published: (2024)