SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911316104445952 |
|---|---|
| author | Shi, Minglei Wang, Haolin Zhang, Borui Zheng, Wenzhao Zeng, Bohan Yuan, Ziyang Wu, Xiaoshi Zhang, Yuanxing Yang, Huan Wang, Xintao Wan, Pengfei Gai, Kun Zhou, Jie Lu, Jiwen |
| author_facet | Shi, Minglei Wang, Haolin Zhang, Borui Zheng, Wenzhao Zeng, Bohan Yuan, Ziyang Wu, Xiaoshi Zhang, Yuanxing Yang, Huan Wang, Xintao Wan, Pengfei Gai, Kun Zhou, Jie Lu, Jiwen |
| contents | Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generation. Despite this potential, training large-scale text-to-image diffusion models entirely within the VFM representation space remains largely unexplored. To bridge this gap, we scale the SVG (Self-supervised representations for Visual Generation) framework, proposing SVG-T2I to support high-quality text-to-image synthesis directly in the VFM feature domain. By leveraging a standard text-to-image diffusion pipeline, SVG-T2I achieves competitive performance, reaching 0.75 on GenEval and 85.78 on DPG-Bench. This performance validates the intrinsic representational power of VFMs for generative tasks. We fully open-source the project, including the autoencoder and generation model, together with their training, inference, evaluation pipelines, and pre-trained weights, to facilitate further research in representation-driven visual generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_11749 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder Shi, Minglei Wang, Haolin Zhang, Borui Zheng, Wenzhao Zeng, Bohan Yuan, Ziyang Wu, Xiaoshi Zhang, Yuanxing Yang, Huan Wang, Xintao Wan, Pengfei Gai, Kun Zhou, Jie Lu, Jiwen Computer Vision and Pattern Recognition Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generation. Despite this potential, training large-scale text-to-image diffusion models entirely within the VFM representation space remains largely unexplored. To bridge this gap, we scale the SVG (Self-supervised representations for Visual Generation) framework, proposing SVG-T2I to support high-quality text-to-image synthesis directly in the VFM feature domain. By leveraging a standard text-to-image diffusion pipeline, SVG-T2I achieves competitive performance, reaching 0.75 on GenEval and 85.78 on DPG-Bench. This performance validates the intrinsic representational power of VFMs for generative tasks. We fully open-source the project, including the autoencoder and generation model, together with their training, inference, evaluation pipelines, and pre-trained weights, to facilitate further research in representation-driven visual generation. |
| title | SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.11749 |