Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pahwa, Ramit, Beedu, Apoorva, Priye, Parivesh, Gandhi, Rutu, Takawale, Saloni, Baijal, Aruna, Yang, Zengli
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914514269634560
author Pahwa, Ramit
Beedu, Apoorva
Priye, Parivesh
Gandhi, Rutu
Takawale, Saloni
Baijal, Aruna
Yang, Zengli
author_facet Pahwa, Ramit
Beedu, Apoorva
Priye, Parivesh
Gandhi, Rutu
Takawale, Saloni
Baijal, Aruna
Yang, Zengli
contents Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning complexity to evaluate tool-calling performance. We introduce Audio2Tool, a large-scale dataset comprising approximately 30,000 queries designed to assess tool-calling capabilities of SpeechLMs across three primary domains: Smart Car, Smart Home, and Wearables. Our benchmark features a multi-tier complexity hierarchy, ranging from simple direct commands to complex multi-intent and needle-in-a-haystack extraction to isolate distinct failure modes. To ensure realism, we employ zero-shot voice cloning text-to-speech synthesis and diverse noise profiles to simulate in-the-wild conditions. Evaluations of state-of-the-art SpeechLMs and ASR-LLM pipelines show strong performance on simple commands but significant degradation under compositional and acoustic challenges. Code and dataset are publicly available on the project page: https://audio2tool.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22821
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
Pahwa, Ramit
Beedu, Apoorva
Priye, Parivesh
Gandhi, Rutu
Takawale, Saloni
Baijal, Aruna
Yang, Zengli
Sound
Machine Learning
Audio and Speech Processing
Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning complexity to evaluate tool-calling performance. We introduce Audio2Tool, a large-scale dataset comprising approximately 30,000 queries designed to assess tool-calling capabilities of SpeechLMs across three primary domains: Smart Car, Smart Home, and Wearables. Our benchmark features a multi-tier complexity hierarchy, ranging from simple direct commands to complex multi-intent and needle-in-a-haystack extraction to isolate distinct failure modes. To ensure realism, we employ zero-shot voice cloning text-to-speech synthesis and diverse noise profiles to simulate in-the-wild conditions. Evaluations of state-of-the-art SpeechLMs and ASR-LLM pipelines show strong performance on simple commands but significant degradation under compositional and acoustic challenges. Code and dataset are publicly available on the project page: https://audio2tool.github.io/.
title Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2604.22821