1.58-bit FLUX

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Chenglin, Liu, Celong, Deng, Xueqing, Kim, Dongwon, Mei, Xing, Shen, Xiaohui, Chen, Liang-Chieh
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913625587843072
author Yang, Chenglin
Liu, Celong
Deng, Xueqing
Kim, Dongwon
Mei, Xing
Shen, Xiaohui
Chen, Liang-Chieh
author_facet Yang, Chenglin
Liu, Celong
Deng, Xueqing
Kim, Dongwon
Mei, Xing
Shen, Xiaohui
Chen, Liang-Chieh
contents We present 1.58-bit FLUX, the first successful approach to quantizing the state-of-the-art text-to-image generation model, FLUX.1-dev, using 1.58-bit weights (i.e., values in {-1, 0, +1}) while maintaining comparable performance for generating 1024 x 1024 images. Notably, our quantization method operates without access to image data, relying solely on self-supervision from the FLUX.1-dev model. Additionally, we develop a custom kernel optimized for 1.58-bit operations, achieving a 7.7x reduction in model storage, a 5.1x reduction in inference memory, and improved inference latency. Extensive evaluations on the GenEval and T2I Compbench benchmarks demonstrate the effectiveness of 1.58-bit FLUX in maintaining generation quality while significantly enhancing computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18653
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 1.58-bit FLUX
Yang, Chenglin
Liu, Celong
Deng, Xueqing
Kim, Dongwon
Mei, Xing
Shen, Xiaohui
Chen, Liang-Chieh
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We present 1.58-bit FLUX, the first successful approach to quantizing the state-of-the-art text-to-image generation model, FLUX.1-dev, using 1.58-bit weights (i.e., values in {-1, 0, +1}) while maintaining comparable performance for generating 1024 x 1024 images. Notably, our quantization method operates without access to image data, relying solely on self-supervision from the FLUX.1-dev model. Additionally, we develop a custom kernel optimized for 1.58-bit operations, achieving a 7.7x reduction in model storage, a 5.1x reduction in inference memory, and improved inference latency. Extensive evaluations on the GenEval and T2I Compbench benchmarks demonstrate the effectiveness of 1.58-bit FLUX in maintaining generation quality while significantly enhancing computational efficiency.
title 1.58-bit FLUX
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.18653