SAR-LM: Symbolic Audio Reasoning with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Taheri, Termeh, Ma, Yinghao, Benetos, Emmanouil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911256275845120
author Taheri, Termeh
Ma, Yinghao
Benetos, Emmanouil
author_facet Taheri, Termeh
Ma, Yinghao
Benetos, Emmanouil
contents Large language models (LLMs) have advanced in text and vision, but their reasoning on audio remains limited. Most existing methods rely on dense audio embeddings, which are difficult to interpret and often fail on structured reasoning tasks. Caption-based approaches, introduced in recent benchmarks such as MMAU, improve performance by translating audio into text, yet still depend on dense embeddings as input, offering little insight when models fail. We present SAR-LM, a symbolic audio reasoning pipeline that builds on this caption-based paradigm by converting audio into structured, human-readable features across speech, sound events, and music. These symbolic inputs support both reasoning and transparent error analysis, enabling us to trace failures to specific features. Across three benchmarks, MMAU, MMAR, and OmniBench, SAR-LM achieves competitive results, while prioritizing interpretability as its primary contribution.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAR-LM: Symbolic Audio Reasoning with Large Language Models
Taheri, Termeh
Ma, Yinghao
Benetos, Emmanouil
Sound
Large language models (LLMs) have advanced in text and vision, but their reasoning on audio remains limited. Most existing methods rely on dense audio embeddings, which are difficult to interpret and often fail on structured reasoning tasks. Caption-based approaches, introduced in recent benchmarks such as MMAU, improve performance by translating audio into text, yet still depend on dense embeddings as input, offering little insight when models fail. We present SAR-LM, a symbolic audio reasoning pipeline that builds on this caption-based paradigm by converting audio into structured, human-readable features across speech, sound events, and music. These symbolic inputs support both reasoning and transparent error analysis, enabling us to trace failures to specific features. Across three benchmarks, MMAU, MMAR, and OmniBench, SAR-LM achieves competitive results, while prioritizing interpretability as its primary contribution.
title SAR-LM: Symbolic Audio Reasoning with Large Language Models
topic Sound
url https://arxiv.org/abs/2511.06483