Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Shilong, Zhang, He, Zhang, Zhifei, Ge, Chongjian, Xue, Shuchen, Liu, Shaoteng, Ren, Mengwei, Kim, Soo Ye, Zhou, Yuqian, Liu, Qing, Pakhomov, Daniil, Zhang, Kai, Lin, Zhe, Luo, Ping
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917158185861120
author Zhang, Shilong
Zhang, He
Zhang, Zhifei
Ge, Chongjian
Xue, Shuchen
Liu, Shaoteng
Ren, Mengwei
Kim, Soo Ye
Zhou, Yuqian
Liu, Qing
Pakhomov, Daniil
Zhang, Kai
Lin, Zhe
Luo, Ping
author_facet Zhang, Shilong
Zhang, He
Zhang, Zhifei
Ge, Chongjian
Xue, Shuchen
Liu, Shaoteng
Ren, Mengwei
Kim, Soo Ye
Zhou, Yuqian
Liu, Qing
Pakhomov, Daniil
Zhang, Kai
Lin, Zhe
Luo, Ping
contents Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation encoders as generative latents. However, we empirically identify two fundamental obstacles in this paradigm: (1) the discriminative feature space lacks compact regularization, making diffusion models prone to off-manifold latents that lead to inaccurate object structures; and (2) the encoder's inherently weak pixel-level reconstruction hinders the generator from learning accurate fine-grained geometry and texture. In this paper, we propose a systematic framework to adapt understanding-oriented encoder features for generative tasks. We introduce a semantic-pixel reconstruction objective to regularize the latent space, enabling the compression of both semantic information and fine-grained details into a highly compact representation (96 channels with 16x16 spatial downsampling). This design ensures that the latent space remains semantically rich and achieves state-of-the-art image reconstruction, while remaining compact enough for accurate generation. Leveraging this representation, we design a unified Text-to-Image (T2I) and image editing model. Benchmarking against various feature spaces, we demonstrate that our approach achieves state-of-the-art reconstruction, faster convergence, and substantial performance gains in both T2I and editing tasks, validating that representation encoders can be effectively adapted into robust generative components.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17909
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
Zhang, Shilong
Zhang, He
Zhang, Zhifei
Ge, Chongjian
Xue, Shuchen
Liu, Shaoteng
Ren, Mengwei
Kim, Soo Ye
Zhou, Yuqian
Liu, Qing
Pakhomov, Daniil
Zhang, Kai
Lin, Zhe
Luo, Ping
Computer Vision and Pattern Recognition
Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation encoders as generative latents. However, we empirically identify two fundamental obstacles in this paradigm: (1) the discriminative feature space lacks compact regularization, making diffusion models prone to off-manifold latents that lead to inaccurate object structures; and (2) the encoder's inherently weak pixel-level reconstruction hinders the generator from learning accurate fine-grained geometry and texture. In this paper, we propose a systematic framework to adapt understanding-oriented encoder features for generative tasks. We introduce a semantic-pixel reconstruction objective to regularize the latent space, enabling the compression of both semantic information and fine-grained details into a highly compact representation (96 channels with 16x16 spatial downsampling). This design ensures that the latent space remains semantically rich and achieves state-of-the-art image reconstruction, while remaining compact enough for accurate generation. Leveraging this representation, we design a unified Text-to-Image (T2I) and image editing model. Benchmarking against various feature spaces, we demonstrate that our approach achieves state-of-the-art reconstruction, faster convergence, and substantial performance gains in both T2I and editing tasks, validating that representation encoders can be effectively adapted into robust generative components.
title Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17909