High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Ji Woo, Yoon, Hee Suk, Koo, Gwanhyeong, Yoon, Eunseop, Eom, SooHwan, Dai, Qi, Luo, Chong, Yoo, Chang D.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914392502697984
author Hong, Ji Woo
Yoon, Hee Suk
Koo, Gwanhyeong
Yoon, Eunseop
Eom, SooHwan
Dai, Qi
Luo, Chong
Yoo, Chang D.
author_facet Hong, Ji Woo
Yoon, Hee Suk
Koo, Gwanhyeong
Yoon, Eunseop
Eom, SooHwan
Dai, Qi
Luo, Chong
Yoo, Chang D.
contents Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image tokenization, which poses a major challenge. Although several studies have explored continuous representation modeling to enhance visual quality, adapting pre-trained VLM models to such representations requires large-scale data and training costs comparable to the original pre-training. To circumvent this limitation, we propose a diffusion-based decoding framework that enhances image fidelity by training only a diffusion decoder on the output image-token logits of pre-trained VLMs, thereby preserving the original model intact. At its core, Logit-to-Code Distributional Mapping converts the VLM's image-token logits into continuous, distribution-weighted code vectors with uncertainty features, providing an effective conditioning signal for diffusion decoding. A lightweight Logit Calibration aligns training-time proxy logits from the VQ-VAE encoder with VLM-generated logits, mitigating the train-inference gap. Conditioned on these representations, the Distribution-Conditioned Diffusion Decoder generates high-fidelity images. Achieved solely through short training on ImageNet-1K, our method consistently improves visual fidelity for both VQ-VAE reconstructions and text-to-image generations from VLM-predicted tokens.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13389
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
Hong, Ji Woo
Yoon, Hee Suk
Koo, Gwanhyeong
Yoon, Eunseop
Eom, SooHwan
Dai, Qi
Luo, Chong
Yoo, Chang D.
Computer Vision and Pattern Recognition
Machine Learning
Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image tokenization, which poses a major challenge. Although several studies have explored continuous representation modeling to enhance visual quality, adapting pre-trained VLM models to such representations requires large-scale data and training costs comparable to the original pre-training. To circumvent this limitation, we propose a diffusion-based decoding framework that enhances image fidelity by training only a diffusion decoder on the output image-token logits of pre-trained VLMs, thereby preserving the original model intact. At its core, Logit-to-Code Distributional Mapping converts the VLM's image-token logits into continuous, distribution-weighted code vectors with uncertainty features, providing an effective conditioning signal for diffusion decoding. A lightweight Logit Calibration aligns training-time proxy logits from the VQ-VAE encoder with VLM-generated logits, mitigating the train-inference gap. Conditioned on these representations, the Distribution-Conditioned Diffusion Decoder generates high-fidelity images. Achieved solely through short training on ImageNet-1K, our method consistently improves visual fidelity for both VQ-VAE reconstructions and text-to-image generations from VLM-predicted tokens.
title High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.13389