Multi-Faceted Evaluation of Tool-Augmented Dialogue Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Zhaoyi Joey, Shourya, Tanya, Wang, Yingfan, Roy, Shamik, Kumar, Vinayshekhar Bannihatti, Gangadharaiah, Rashmi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915569702273024
author Hou, Zhaoyi Joey
Shourya, Tanya
Wang, Yingfan
Roy, Shamik
Kumar, Vinayshekhar Bannihatti
Gangadharaiah, Rashmi
author_facet Hou, Zhaoyi Joey
Shourya, Tanya
Wang, Yingfan
Roy, Shamik
Kumar, Vinayshekhar Bannihatti
Gangadharaiah, Rashmi
contents Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and tools. While existing evaluation methods assess either user satisfaction or agents' tool-calling capabilities, they fail to capture critical errors in multi-turn tool-augmented dialogues-such as when agents misinterpret tool results yet appear satisfactory to users. We introduce TRACE, a benchmark of systematically synthesized tool-augmented conversations covering diverse error cases, and SCOPE, an evaluation framework that automatically discovers diverse error patterns and evaluation rubrics in tool-augmented dialogues. Experiments show SCOPE significantly outperforms the baseline, particularly on challenging cases where user satisfaction signals are misleading.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Faceted Evaluation of Tool-Augmented Dialogue Systems
Hou, Zhaoyi Joey
Shourya, Tanya
Wang, Yingfan
Roy, Shamik
Kumar, Vinayshekhar Bannihatti
Gangadharaiah, Rashmi
Computation and Language
Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and tools. While existing evaluation methods assess either user satisfaction or agents' tool-calling capabilities, they fail to capture critical errors in multi-turn tool-augmented dialogues-such as when agents misinterpret tool results yet appear satisfactory to users. We introduce TRACE, a benchmark of systematically synthesized tool-augmented conversations covering diverse error cases, and SCOPE, an evaluation framework that automatically discovers diverse error patterns and evaluation rubrics in tool-augmented dialogues. Experiments show SCOPE significantly outperforms the baseline, particularly on challenging cases where user satisfaction signals are misleading.
title Multi-Faceted Evaluation of Tool-Augmented Dialogue Systems
topic Computation and Language
url https://arxiv.org/abs/2510.19186