ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sienkiewicz, Bruno, Neumann, Łukasz, Modrzejewski, Mateusz
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915772302884864
author Sienkiewicz, Bruno
Neumann, Łukasz
Modrzejewski, Mateusz
author_facet Sienkiewicz, Bruno
Neumann, Łukasz
Modrzejewski, Mateusz
contents Concept-based interpretability methods like TCAV require clean, well-separated positive and negative examples for each concept. Existing music datasets lack this structure: tags are sparse, noisy, or ill-defined. We introduce ConceptCaps, a dataset of 21k music-caption-tags triplets with explicit labels from a 200-attribute taxonomy. Our pipeline separates semantic modeling from text generation: a VAE learns plausible attribute co-occurrence patterns, a fine-tuned LLM converts attribute lists into professional descriptions, and MusicGen synthesizes corresponding audio. This separation improves coherence and controllability over end-to-end approaches. We validate the dataset through audio-text alignment (CLAP), linguistic quality metrics (BERTScore, MAUVE), and TCAV analysis confirming that concept probes recover musically meaningful patterns. Dataset and code are available online.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14157
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models
Sienkiewicz, Bruno
Neumann, Łukasz
Modrzejewski, Mateusz
Sound
Artificial Intelligence
Machine Learning
Concept-based interpretability methods like TCAV require clean, well-separated positive and negative examples for each concept. Existing music datasets lack this structure: tags are sparse, noisy, or ill-defined. We introduce ConceptCaps, a dataset of 21k music-caption-tags triplets with explicit labels from a 200-attribute taxonomy. Our pipeline separates semantic modeling from text generation: a VAE learns plausible attribute co-occurrence patterns, a fine-tuned LLM converts attribute lists into professional descriptions, and MusicGen synthesizes corresponding audio. This separation improves coherence and controllability over end-to-end approaches. We validate the dataset through audio-text alignment (CLAP), linguistic quality metrics (BERTScore, MAUVE), and TCAV analysis confirming that concept probes recover musically meaningful patterns. Dataset and code are available online.
title ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models
topic Sound
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.14157