DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wilkinghoff, Kevin, Tan, Zheng-Hua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911245596098560
author Wilkinghoff, Kevin
Tan, Zheng-Hua
author_facet Wilkinghoff, Kevin
Tan, Zheng-Hua
contents Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the type of sound events, as well as the direction and distance of their corresponding sources. Accomplishing this with a single audio encoder is demanding as the information required for each of these tasks is mostly independent of each other. As a result, the performance obtained with a single encoder is often worse than when using task-specific audio encoders. In this work, we present DSpAST, a novel audio encoder based on SpatialAST that learns disentangled representations of spatial audio while having only 0.2% additional parameters. Experiments on SpatialSoundQA with the spatial audio reasoning system BAT demonstrate that DSpAST significantly outperforms SpatialAST.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13927
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
Wilkinghoff, Kevin
Tan, Zheng-Hua
Audio and Speech Processing
Artificial Intelligence
Sound
Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the type of sound events, as well as the direction and distance of their corresponding sources. Accomplishing this with a single audio encoder is demanding as the information required for each of these tasks is mostly independent of each other. As a result, the performance obtained with a single encoder is often worse than when using task-specific audio encoders. In this work, we present DSpAST, a novel audio encoder based on SpatialAST that learns disentangled representations of spatial audio while having only 0.2% additional parameters. Experiments on SpatialSoundQA with the spatial audio reasoning system BAT demonstrate that DSpAST significantly outperforms SpatialAST.
title DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2509.13927