Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Kuan-Yi, Lin, Tsung-En, Lee, Hung-Yi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912645064425472
author Lee, Kuan-Yi
Lin, Tsung-En
Lee, Hung-Yi
author_facet Lee, Kuan-Yi
Lin, Tsung-En
Lee, Hung-Yi
contents Recent advancements in large multimodal models (LMMs) have shown strong capabilities in audio understanding. However, most systems rely solely on end-to-end reasoning, limiting interpretability and accuracy for tasks that require structured knowledge or specialized signal analysis. In this work, we present Audio-Maestro -- a tool-augmented audio reasoning framework that enables audio-language models to autonomously call external tools and integrate their timestamped outputs into the reasoning process. This design allows the model to analyze, transform, and interpret audio signals through specialized tools rather than relying solely on end-to-end inference. Experiments show that Audio-Maestro consistently improves general audio reasoning performance: Gemini-2.5-flash's average accuracy on MMAU-Test rises from 67.4% to 72.1%, DeSTA-2.5 from 58.3% to 62.8%, and GPT-4o from 60.8% to 63.9%. To our knowledge, Audio-Maestro is the first framework to integrate structured tool output into the large audio language model reasoning process.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
Lee, Kuan-Yi
Lin, Tsung-En
Lee, Hung-Yi
Sound
Artificial Intelligence
Recent advancements in large multimodal models (LMMs) have shown strong capabilities in audio understanding. However, most systems rely solely on end-to-end reasoning, limiting interpretability and accuracy for tasks that require structured knowledge or specialized signal analysis. In this work, we present Audio-Maestro -- a tool-augmented audio reasoning framework that enables audio-language models to autonomously call external tools and integrate their timestamped outputs into the reasoning process. This design allows the model to analyze, transform, and interpret audio signals through specialized tools rather than relying solely on end-to-end inference. Experiments show that Audio-Maestro consistently improves general audio reasoning performance: Gemini-2.5-flash's average accuracy on MMAU-Test rises from 67.4% to 72.1%, DeSTA-2.5 from 58.3% to 62.8%, and GPT-4o from 60.8% to 63.9%. To our knowledge, Audio-Maestro is the first framework to integrate structured tool output into the large audio language model reasoning process.
title Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2510.11454