FraQAT: Quantization Aware Training with Fractional bits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Morreale, Luca, Ramos, Alberto Gil C. P., Chadwick, Malcolm, Noroozi, Mehid, Chavhan, Ruchika, Mehrotra, Abhinav, Bhattacharya, Sourav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909850347241472
author Morreale, Luca
Ramos, Alberto Gil C. P.
Chadwick, Malcolm
Noroozi, Mehid
Chavhan, Ruchika
Mehrotra, Abhinav
Bhattacharya, Sourav
author_facet Morreale, Luca
Ramos, Alberto Gil C. P.
Chadwick, Malcolm
Noroozi, Mehid
Chavhan, Ruchika
Mehrotra, Abhinav
Bhattacharya, Sourav
contents State-of-the-art (SOTA) generative models have demonstrated impressive capabilities in image synthesis or text generation, often with a large capacity model. However, these large models cannot be deployed on smartphones due to the limited availability of on-board memory and computations. Quantization methods lower the precision of the model parameters, allowing for efficient computations, \eg, in \INT{8}. Although aggressive quantization addresses efficiency and memory constraints, preserving the quality of the model remains a challenge. To retain quality in previous aggressive quantization, we propose a new fractional bits quantization (\short) approach. The novelty is a simple yet effective idea: we progressively reduce the model's precision from 32 to 4 bits per parameter, and exploit the fractional bits during optimization to maintain high generation quality. We show that the \short{} yields improved quality on a variety of diffusion models, including SD3.5-Medium, Sana, \pixart, and FLUX.1-schnell, while achieving $4-7\%$ lower FiD than standard QAT. Finally, we deploy and run Sana on a Samsung S25U, which runs on the Qualcomm SM8750-AB Snapdragon 8 Elite Hexagon Tensor Processor (HTP).
format Preprint
id arxiv_https___arxiv_org_abs_2510_14823
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FraQAT: Quantization Aware Training with Fractional bits
Morreale, Luca
Ramos, Alberto Gil C. P.
Chadwick, Malcolm
Noroozi, Mehid
Chavhan, Ruchika
Mehrotra, Abhinav
Bhattacharya, Sourav
Computer Vision and Pattern Recognition
State-of-the-art (SOTA) generative models have demonstrated impressive capabilities in image synthesis or text generation, often with a large capacity model. However, these large models cannot be deployed on smartphones due to the limited availability of on-board memory and computations. Quantization methods lower the precision of the model parameters, allowing for efficient computations, \eg, in \INT{8}. Although aggressive quantization addresses efficiency and memory constraints, preserving the quality of the model remains a challenge. To retain quality in previous aggressive quantization, we propose a new fractional bits quantization (\short) approach. The novelty is a simple yet effective idea: we progressively reduce the model's precision from 32 to 4 bits per parameter, and exploit the fractional bits during optimization to maintain high generation quality. We show that the \short{} yields improved quality on a variety of diffusion models, including SD3.5-Medium, Sana, \pixart, and FLUX.1-schnell, while achieving $4-7\%$ lower FiD than standard QAT. Finally, we deploy and run Sana on a Samsung S25U, which runs on the Qualcomm SM8750-AB Snapdragon 8 Elite Hexagon Tensor Processor (HTP).
title FraQAT: Quantization Aware Training with Fractional bits
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14823