Low-Resource Guidance for Controllable Latent Audio Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Novack, Zachary, Zukowski, Zack, Carr, CJ, Parker, Julian, Evans, Zach, Taylor, Josiah, Berg-Kirkpatrick, Taylor, McAuley, Julian, Pons, Jordi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912942780317696
author Novack, Zachary
Zukowski, Zack
Carr, CJ
Parker, Julian
Evans, Zach
Taylor, Josiah
Berg-Kirkpatrick, Taylor
McAuley, Julian
Pons, Jordi
author_facet Novack, Zachary
Zukowski, Zack
Carr, CJ
Parker, Julian
Evans, Zach
Taylor, Josiah
Berg-Kirkpatrick, Taylor
McAuley, Julian
Pons, Jordi
contents Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guidance) that can also be computationally demanding. By examining the bottlenecks of existing guidance-based controls, in particular their high cost-per-step due to decoder backpropagation, we introduce a guidance-based approach through selective TFG and Latent-Control Heads (LatCHs), which enables controlling latent audio diffusion models with low computational overhead. LatCHs operate directly in latent space, avoiding the expensive decoder step, and requiring minimal training resources (7M parameters and $\approx$ 4 hours of training). Experiments with Stable Audio Open demonstrate effective control over intensity, pitch, and beats (and a combination of those) while maintaining generation quality. Our method balances precision and audio fidelity with far lower computational costs than standard end-to-end guidance. Demo examples can be found at https://zacharynovack.github.io/latch/latch.html.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Low-Resource Guidance for Controllable Latent Audio Diffusion
Novack, Zachary
Zukowski, Zack
Carr, CJ
Parker, Julian
Evans, Zach
Taylor, Josiah
Berg-Kirkpatrick, Taylor
McAuley, Julian
Pons, Jordi
Sound
Artificial Intelligence
Machine Learning
Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guidance) that can also be computationally demanding. By examining the bottlenecks of existing guidance-based controls, in particular their high cost-per-step due to decoder backpropagation, we introduce a guidance-based approach through selective TFG and Latent-Control Heads (LatCHs), which enables controlling latent audio diffusion models with low computational overhead. LatCHs operate directly in latent space, avoiding the expensive decoder step, and requiring minimal training resources (7M parameters and $\approx$ 4 hours of training). Experiments with Stable Audio Open demonstrate effective control over intensity, pitch, and beats (and a combination of those) while maintaining generation quality. Our method balances precision and audio fidelity with far lower computational costs than standard end-to-end guidance. Demo examples can be found at https://zacharynovack.github.io/latch/latch.html.
title Low-Resource Guidance for Controllable Latent Audio Diffusion
topic Sound
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.04366