SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Koo, Junghyun, Wichern, Gordon, Germain, Francois G., Khurana, Sameer, Roux, Jonathan Le
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917907909312512
author Koo, Junghyun
Wichern, Gordon
Germain, Francois G.
Khurana, Sameer
Roux, Jonathan Le
author_facet Koo, Junghyun
Wichern, Gordon
Germain, Francois G.
Khurana, Sameer
Roux, Jonathan Le
contents We introduce Self-Monitored Inference-Time INtervention (SMITIN), an approach for controlling an autoregressive generative music transformer using classifier probes. These simple logistic regression probes are trained on the output of each attention head in the transformer using a small dataset of audio examples both exhibiting and missing a specific musical trait (e.g., the presence/absence of drums, or real/synthetic music). We then steer the attention heads in the probe direction, ensuring the generative model output captures the desired musical trait. Additionally, we monitor the probe output to avoid adding an excessive amount of intervention into the autoregressive generation, which could lead to temporally incoherent music. We validate our results objectively and subjectively for both audio continuation and text-to-music applications, demonstrating the ability to add controls to large generative models for which retraining or even fine-tuning is impractical for most musicians. Audio samples of the proposed intervention approach are available on our demo page http://tinyurl.com/smitin .
format Preprint
id arxiv_https___arxiv_org_abs_2404_02252
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers
Koo, Junghyun
Wichern, Gordon
Germain, Francois G.
Khurana, Sameer
Roux, Jonathan Le
Sound
Audio and Speech Processing
We introduce Self-Monitored Inference-Time INtervention (SMITIN), an approach for controlling an autoregressive generative music transformer using classifier probes. These simple logistic regression probes are trained on the output of each attention head in the transformer using a small dataset of audio examples both exhibiting and missing a specific musical trait (e.g., the presence/absence of drums, or real/synthetic music). We then steer the attention heads in the probe direction, ensuring the generative model output captures the desired musical trait. Additionally, we monitor the probe output to avoid adding an excessive amount of intervention into the autoregressive generation, which could lead to temporally incoherent music. We validate our results objectively and subjectively for both audio continuation and text-to-music applications, demonstrating the ability to add controls to large generative models for which retraining or even fine-tuning is impractical for most musicians. Audio samples of the proposed intervention approach are available on our demo page http://tinyurl.com/smitin .
title SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2404.02252