LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Xianfeng, Bai, Yajing, Zheng, Haoze, Chen, Harold Haodong, Liu, Yexin, Wang, Zihao, Ma, Xuran, Shu, Wen-Jie, Wu, Xianzu, Yang, Harry, Lim, Ser-Nam
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909534559141888
author Wu, Xianfeng
Bai, Yajing
Zheng, Haoze
Chen, Harold Haodong
Liu, Yexin
Wang, Zihao
Ma, Xuran
Shu, Wen-Jie
Wu, Xianzu
Yang, Harry
Lim, Ser-Nam
author_facet Wu, Xianfeng
Bai, Yajing
Zheng, Haoze
Chen, Harold Haodong
Liu, Yexin
Wang, Zihao
Ma, Xuran
Shu, Wen-Jie
Wu, Xianzu
Yang, Harry
Lim, Ser-Nam
contents Recent advances in text-to-image generation have primarily relied on extensive datasets and parameter-heavy architectures. These requirements severely limit accessibility for researchers and practitioners who lack substantial computational resources. In this paper, we introduce \model, an efficient training paradigm for image generation models that uses knowledge distillation (KD) and Direct Preference Optimization (DPO). Drawing inspiration from the success of data KD techniques widely adopted in Multi-Modal Large Language Models (MLLMs), LightGen distills knowledge from state-of-the-art (SOTA) text-to-image models into a compact Masked Autoregressive (MAR) architecture with only $0.7B$ parameters. Using a compact synthetic dataset of just $2M$ high-quality images generated from varied captions, we demonstrate that data diversity significantly outweighs data volume in determining model performance. This strategy dramatically reduces computational demands and reduces pre-training time from potentially thousands of GPU-days to merely 88 GPU-days. Furthermore, to address the inherent shortcomings of synthetic data, particularly poor high-frequency details and spatial inaccuracies, we integrate the DPO technique that refines image fidelity and positional accuracy. Comprehensive experiments confirm that LightGen achieves image generation quality comparable to SOTA models while significantly reducing computational resources and expanding accessibility for resource-constrained environments. Code is available at https://github.com/XianfengWu01/LightGen
format Preprint
id arxiv_https___arxiv_org_abs_2503_08619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization
Wu, Xianfeng
Bai, Yajing
Zheng, Haoze
Chen, Harold Haodong
Liu, Yexin
Wang, Zihao
Ma, Xuran
Shu, Wen-Jie
Wu, Xianzu
Yang, Harry
Lim, Ser-Nam
Computer Vision and Pattern Recognition
Recent advances in text-to-image generation have primarily relied on extensive datasets and parameter-heavy architectures. These requirements severely limit accessibility for researchers and practitioners who lack substantial computational resources. In this paper, we introduce \model, an efficient training paradigm for image generation models that uses knowledge distillation (KD) and Direct Preference Optimization (DPO). Drawing inspiration from the success of data KD techniques widely adopted in Multi-Modal Large Language Models (MLLMs), LightGen distills knowledge from state-of-the-art (SOTA) text-to-image models into a compact Masked Autoregressive (MAR) architecture with only $0.7B$ parameters. Using a compact synthetic dataset of just $2M$ high-quality images generated from varied captions, we demonstrate that data diversity significantly outweighs data volume in determining model performance. This strategy dramatically reduces computational demands and reduces pre-training time from potentially thousands of GPU-days to merely 88 GPU-days. Furthermore, to address the inherent shortcomings of synthetic data, particularly poor high-frequency details and spatial inaccuracies, we integrate the DPO technique that refines image fidelity and positional accuracy. Comprehensive experiments confirm that LightGen achieves image generation quality comparable to SOTA models while significantly reducing computational resources and expanding accessibility for resource-constrained environments. Code is available at https://github.com/XianfengWu01/LightGen
title LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.08619