PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vora, Jayneel, Krishnan, Aditya, Bouacida, Nader, Shankar, Prabhu RV, Mohapatra, Prasant
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913512006090752
author Vora, Jayneel
Krishnan, Aditya
Bouacida, Nader
Shankar, Prabhu RV
Mohapatra, Prasant
author_facet Vora, Jayneel
Krishnan, Aditya
Bouacida, Nader
Shankar, Prabhu RV
Mohapatra, Prasant
contents Denoising diffusion models have emerged as state-of-the-art in generative tasks across image, audio, and video domains, producing high-quality, diverse, and contextually relevant data. However, their broader adoption is limited by high computational costs and large memory footprints. Post-training quantization (PTQ) offers a promising approach to mitigate these challenges by reducing model complexity through low-bandwidth parameters. Yet, direct application of PTQ to diffusion models can degrade synthesis quality due to accumulated quantization noise across multiple denoising steps, particularly in conditional tasks like text-to-audio synthesis. This work introduces PTQ4ADM, a novel framework for quantizing audio diffusion models(ADMs). Our key contributions include (1) a coverage-driven prompt augmentation method and (2) an activation-aware calibration set generation algorithm for text-conditional ADMs. These techniques ensure comprehensive coverage of audio aspects and modalities while preserving synthesis fidelity. We validate our approach on TANGO, Make-An-Audio, and AudioLDM models for text-conditional audio generation. Extensive experiments demonstrate PTQ4ADM's capability to reduce the model size by up to 70\% while achieving synthesis quality metrics comparable to full-precision models($<$5\% increase in FD scores). We show that specific layers in the backbone network can be quantized to 4-bit weights and 8-bit activations without significant quality loss. This work paves the way for more efficient deployment of ADMs in resource-constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13894
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
Vora, Jayneel
Krishnan, Aditya
Bouacida, Nader
Shankar, Prabhu RV
Mohapatra, Prasant
Sound
Machine Learning
Audio and Speech Processing
Denoising diffusion models have emerged as state-of-the-art in generative tasks across image, audio, and video domains, producing high-quality, diverse, and contextually relevant data. However, their broader adoption is limited by high computational costs and large memory footprints. Post-training quantization (PTQ) offers a promising approach to mitigate these challenges by reducing model complexity through low-bandwidth parameters. Yet, direct application of PTQ to diffusion models can degrade synthesis quality due to accumulated quantization noise across multiple denoising steps, particularly in conditional tasks like text-to-audio synthesis. This work introduces PTQ4ADM, a novel framework for quantizing audio diffusion models(ADMs). Our key contributions include (1) a coverage-driven prompt augmentation method and (2) an activation-aware calibration set generation algorithm for text-conditional ADMs. These techniques ensure comprehensive coverage of audio aspects and modalities while preserving synthesis fidelity. We validate our approach on TANGO, Make-An-Audio, and AudioLDM models for text-conditional audio generation. Extensive experiments demonstrate PTQ4ADM's capability to reduce the model size by up to 70\% while achieving synthesis quality metrics comparable to full-precision models($<$5\% increase in FD scores). We show that specific layers in the backbone network can be quantized to 4-bit weights and 8-bit activations without significant quality loss. This work paves the way for more efficient deployment of ADMs in resource-constrained environments.
title PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2409.13894