Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Ze, Chen, Hao, Hu, Benran, Liu, Jiang, Sun, Ximeng, Wu, Jialian, Su, Yusheng, Yu, Xiaodong, Barsoum, Emad, Liu, Zicheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916811473158144
author Wang, Ze
Chen, Hao
Hu, Benran
Liu, Jiang
Sun, Ximeng
Wu, Jialian
Su, Yusheng
Yu, Xiaodong
Barsoum, Emad
Liu, Zicheng
author_facet Wang, Ze
Chen, Hao
Hu, Benran
Liu, Jiang
Sun, Ximeng
Wu, Jialian
Su, Yusheng
Yu, Xiaodong
Barsoum, Emad
Liu, Zicheng
contents Image tokenization plays a critical role in reducing the computational demands of modeling high-resolution images, significantly improving the efficiency of image and multimodal understanding and generation. Recent advances in 1D latent spaces have reduced the number of tokens required by eliminating the need for a 2D grid structure. In this paper, we further advance compact discrete image representation by introducing 1D binary image latents. By representing each image as a sequence of binary vectors, rather than using traditional one-hot codebook tokens, our approach preserves high-resolution details while maintaining the compactness of 1D latents. To the best of our knowledge, our text-to-image models are the first to achieve competitive performance in both diffusion and auto-regressive generation using just 128 discrete tokens for images up to 1024x1024, demonstrating up to a 32-fold reduction in token numbers compared to standard VQ-VAEs. The proposed 1D binary latent space, coupled with simple model architectures, achieves marked improvements in speed training and inference speed. Our text-to-image models allow for a global batch size of 4096 on a single GPU node with 8 AMD MI300X GPUs, and the training can be completed within 200 GPU days. Our models achieve competitive performance compared to modern image generation models without any in-house private training data or post-training refinements, offering a scalable and efficient alternative to conventional tokenization methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
Wang, Ze
Chen, Hao
Hu, Benran
Liu, Jiang
Sun, Ximeng
Wu, Jialian
Su, Yusheng
Yu, Xiaodong
Barsoum, Emad
Liu, Zicheng
Computer Vision and Pattern Recognition
Image tokenization plays a critical role in reducing the computational demands of modeling high-resolution images, significantly improving the efficiency of image and multimodal understanding and generation. Recent advances in 1D latent spaces have reduced the number of tokens required by eliminating the need for a 2D grid structure. In this paper, we further advance compact discrete image representation by introducing 1D binary image latents. By representing each image as a sequence of binary vectors, rather than using traditional one-hot codebook tokens, our approach preserves high-resolution details while maintaining the compactness of 1D latents. To the best of our knowledge, our text-to-image models are the first to achieve competitive performance in both diffusion and auto-regressive generation using just 128 discrete tokens for images up to 1024x1024, demonstrating up to a 32-fold reduction in token numbers compared to standard VQ-VAEs. The proposed 1D binary latent space, coupled with simple model architectures, achieves marked improvements in speed training and inference speed. Our text-to-image models allow for a global batch size of 4096 on a single GPU node with 8 AMD MI300X GPUs, and the training can be completed within 200 GPU days. Our models achieve competitive performance compared to modern image generation models without any in-house private training data or post-training refinements, offering a scalable and efficient alternative to conventional tokenization methods.
title Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21022