ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khandelwal, Dinesh, Punnavajhala, Gnana Prakash, Bhargav, GPS, Pandey, Gaurav, Joshi, Sachin, Karanam, Hima, Raghu, Dinesh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911676840804352
author Khandelwal, Dinesh
Punnavajhala, Gnana Prakash
Bhargav, GPS
Pandey, Gaurav
Joshi, Sachin
Karanam, Hima
Raghu, Dinesh
author_facet Khandelwal, Dinesh
Punnavajhala, Gnana Prakash
Bhargav, GPS
Pandey, Gaurav
Joshi, Sachin
Karanam, Hima
Raghu, Dinesh
contents Multi-turn tool calling is essential for LLMs to function as autonomous agents, yet synthesizing the training data required for these capabilities remains a fundamental challenge. Existing synthetic data generation pipelines often produce unrealistic dialogues for two reasons: they chain tools that are only superficially compatible rather than aligned with meaningful user tasks, and they generate dialogues in one shot, which often introduces arguments that were neither provided by the user nor produced by prior tool calls. These issues also lead to a severe underrepresentation of multi-step tool interactions. We introduce ToolWeave, a structured framework for synthesizing realistic multi-turn tool-calling dialogues. ToolWeave support realistic multi-step workflows (or tool sequences) by constructing tools with built-in dependencies and filters the workflows based on alignment with user goals. It reduces parameter hallucination by using a fine-grained planning stage that explicitly tracks parameter provenance. As a result, ToolWeave-generated synthetic dialogues contain more multi-step tool interactions (45%) and fewer hallucinations in parameters and tool names. Consequently, LLMs fine-tuned on ToolWeave consistently outperform those fine-tuned on prior datasets across three public benchmarks. Notably, Llama-3.1-70B fine-tuned on ToolWeave achieves 39.75% on BFCL-V3 multi-turn, compared to 23.50% when fine-tuned on SOTA ToolFlow data.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12521
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues
Khandelwal, Dinesh
Punnavajhala, Gnana Prakash
Bhargav, GPS
Pandey, Gaurav
Joshi, Sachin
Karanam, Hima
Raghu, Dinesh
Computation and Language
Artificial Intelligence
Multi-turn tool calling is essential for LLMs to function as autonomous agents, yet synthesizing the training data required for these capabilities remains a fundamental challenge. Existing synthetic data generation pipelines often produce unrealistic dialogues for two reasons: they chain tools that are only superficially compatible rather than aligned with meaningful user tasks, and they generate dialogues in one shot, which often introduces arguments that were neither provided by the user nor produced by prior tool calls. These issues also lead to a severe underrepresentation of multi-step tool interactions. We introduce ToolWeave, a structured framework for synthesizing realistic multi-turn tool-calling dialogues. ToolWeave support realistic multi-step workflows (or tool sequences) by constructing tools with built-in dependencies and filters the workflows based on alignment with user goals. It reduces parameter hallucination by using a fine-grained planning stage that explicitly tracks parameter provenance. As a result, ToolWeave-generated synthetic dialogues contain more multi-step tool interactions (45%) and fewer hallucinations in parameters and tool names. Consequently, LLMs fine-tuned on ToolWeave consistently outperform those fine-tuned on prior datasets across three public benchmarks. Notably, Llama-3.1-70B fine-tuned on ToolWeave achieves 39.75% on BFCL-V3 multi-turn, compared to 23.50% when fine-tuned on SOTA ToolFlow data.
title ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.12521