ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shim, Jeonghoon, Seo, Gyuhyeon, Lim, Cheongsu, Jo, Yohan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912255203868672
author Shim, Jeonghoon
Seo, Gyuhyeon
Lim, Cheongsu
Jo, Yohan
author_facet Shim, Jeonghoon
Seo, Gyuhyeon
Lim, Cheongsu
Jo, Yohan
contents Tool-Augmented Language Models (TALMs) leverage external APIs to answer user queries across various domains. However, existing benchmark datasets for TALM research often feature simplistic dialogues that do not reflect real-world scenarios, such as the need for models to ask clarifying questions or proactively call additional APIs when essential information is missing. To address these limitations, we construct and release ToolDial, a dataset comprising 11,111 multi-turn dialogues, with an average of 8.95 turns per dialogue, based on APIs from RapidAPI. ToolDial has two key characteristics. First, the dialogues incorporate 16 user and system actions (e.g., "Request", "Clarify", "Fail inform") to capture the rich dynamics of real-world interactions. Second, we simulate dialogues where the system requests necessary information from the user based on API documentation and seeks additional APIs if the user fails to provide the required information. To facilitate this process, we introduce a method for generating an API graph that represents input and output compatibility between APIs. Using ToolDial, we evaluate a suite of language models on their ability to predict correct actions and extract input parameter values for API calls from the dialogue history. Modern language models achieve accuracy scores below 70%, indicating substantial room for improvement. We release our dataset and code at https://github.com/holi-lab/ToolDial.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00564
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
Shim, Jeonghoon
Seo, Gyuhyeon
Lim, Cheongsu
Jo, Yohan
Computation and Language
Tool-Augmented Language Models (TALMs) leverage external APIs to answer user queries across various domains. However, existing benchmark datasets for TALM research often feature simplistic dialogues that do not reflect real-world scenarios, such as the need for models to ask clarifying questions or proactively call additional APIs when essential information is missing. To address these limitations, we construct and release ToolDial, a dataset comprising 11,111 multi-turn dialogues, with an average of 8.95 turns per dialogue, based on APIs from RapidAPI. ToolDial has two key characteristics. First, the dialogues incorporate 16 user and system actions (e.g., "Request", "Clarify", "Fail inform") to capture the rich dynamics of real-world interactions. Second, we simulate dialogues where the system requests necessary information from the user based on API documentation and seeks additional APIs if the user fails to provide the required information. To facilitate this process, we introduce a method for generating an API graph that represents input and output compatibility between APIs. Using ToolDial, we evaluate a suite of language models on their ability to predict correct actions and extract input parameter values for API calls from the dialogue history. Modern language models achieve accuracy scores below 70%, indicating substantial room for improvement. We release our dataset and code at https://github.com/holi-lab/ToolDial.
title ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
topic Computation and Language
url https://arxiv.org/abs/2503.00564