Evaluating Compositional Structure in Audio Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Chuyang, Steers, Bea, McFee, Brian, Bello, Juan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914392908496896
author Chen, Chuyang
Steers, Bea
McFee, Brian
Bello, Juan
author_facet Chen, Chuyang
Steers, Bea
McFee, Brian
Bello, Juan
contents We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attributes, and combining them systematically. While central to auditory perception, this property is largely absent from current evaluation protocols. Our framework adapts ideas from vision and language to audio through two tasks: A-COAT, which tests consistency under additive transformations, and A-TRE, which probes reconstructibility from attribute-level primitives. Both tasks are supported by large synthetic datasets with controlled variation in acoustic attributes, providing the first benchmark of compositional structure in audio embeddings.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13685
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating Compositional Structure in Audio Representations
Chen, Chuyang
Steers, Bea
McFee, Brian
Bello, Juan
Sound
We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attributes, and combining them systematically. While central to auditory perception, this property is largely absent from current evaluation protocols. Our framework adapts ideas from vision and language to audio through two tasks: A-COAT, which tests consistency under additive transformations, and A-TRE, which probes reconstructibility from attribute-level primitives. Both tasks are supported by large synthetic datasets with controlled variation in acoustic attributes, providing the first benchmark of compositional structure in audio embeddings.
title Evaluating Compositional Structure in Audio Representations
topic Sound
url https://arxiv.org/abs/2603.13685