SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Minglei, Wang, Haolin, Zhang, Borui, Zheng, Wenzhao, Zeng, Bohan, Yuan, Ziyang, Wu, Xiaoshi, Zhang, Yuanxing, Yang, Huan, Wang, Xintao, Wan, Pengfei, Gai, Kun, Zhou, Jie, Lu, Jiwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911316104445952
author Shi, Minglei
Wang, Haolin
Zhang, Borui
Zheng, Wenzhao
Zeng, Bohan
Yuan, Ziyang
Wu, Xiaoshi
Zhang, Yuanxing
Yang, Huan
Wang, Xintao
Wan, Pengfei
Gai, Kun
Zhou, Jie
Lu, Jiwen
author_facet Shi, Minglei
Wang, Haolin
Zhang, Borui
Zheng, Wenzhao
Zeng, Bohan
Yuan, Ziyang
Wu, Xiaoshi
Zhang, Yuanxing
Yang, Huan
Wang, Xintao
Wan, Pengfei
Gai, Kun
Zhou, Jie
Lu, Jiwen
contents Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generation. Despite this potential, training large-scale text-to-image diffusion models entirely within the VFM representation space remains largely unexplored. To bridge this gap, we scale the SVG (Self-supervised representations for Visual Generation) framework, proposing SVG-T2I to support high-quality text-to-image synthesis directly in the VFM feature domain. By leveraging a standard text-to-image diffusion pipeline, SVG-T2I achieves competitive performance, reaching 0.75 on GenEval and 85.78 on DPG-Bench. This performance validates the intrinsic representational power of VFMs for generative tasks. We fully open-source the project, including the autoencoder and generation model, together with their training, inference, evaluation pipelines, and pre-trained weights, to facilitate further research in representation-driven visual generation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
Shi, Minglei
Wang, Haolin
Zhang, Borui
Zheng, Wenzhao
Zeng, Bohan
Yuan, Ziyang
Wu, Xiaoshi
Zhang, Yuanxing
Yang, Huan
Wang, Xintao
Wan, Pengfei
Gai, Kun
Zhou, Jie
Lu, Jiwen
Computer Vision and Pattern Recognition
Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generation. Despite this potential, training large-scale text-to-image diffusion models entirely within the VFM representation space remains largely unexplored. To bridge this gap, we scale the SVG (Self-supervised representations for Visual Generation) framework, proposing SVG-T2I to support high-quality text-to-image synthesis directly in the VFM feature domain. By leveraging a standard text-to-image diffusion pipeline, SVG-T2I achieves competitive performance, reaching 0.75 on GenEval and 85.78 on DPG-Bench. This performance validates the intrinsic representational power of VFMs for generative tasks. We fully open-source the project, including the autoencoder and generation model, together with their training, inference, evaluation pipelines, and pre-trained weights, to facilitate further research in representation-driven visual generation.
title SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11749