REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Leng, Xingjian, Singh, Jaskirat, Hou, Yunzhong, Xing, Zhenchang, Xie, Saining, Zheng, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What matters for Representation Alignment: Global Information or Spatial Structure?
by: Singh, Jaskirat, et al.
Published: (2025)
by: Singh, Jaskirat, et al.
Published: (2025)
SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
by: Zhao, Qinyu, et al.
Published: (2025)
by: Zhao, Qinyu, et al.
Published: (2025)
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025)
by: Zheng, Boyang, et al.
Published: (2025)
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026)
by: Duggal, Shivam, et al.
Published: (2026)
Improved Baselines with Representation Autoencoders
by: Singh, Jaskirat, et al.
Published: (2026)
by: Singh, Jaskirat, et al.
Published: (2026)
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
by: Bi, Tianci, et al.
Published: (2025)
by: Bi, Tianci, et al.
Published: (2025)
Diffusion As Self-Distillation: End-to-End Latent Diffusion In One Model
by: Wang, Xiyuan, et al.
Published: (2025)
by: Wang, Xiyuan, et al.
Published: (2025)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)
by: Lei, Jiachen, et al.
Published: (2025)
SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation
by: Xing, Ximing, et al.
Published: (2024)
by: Xing, Ximing, et al.
Published: (2024)
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
End-to-End Multi-Modal Diffusion Mamba
by: Lu, Chunhao, et al.
Published: (2025)
by: Lu, Chunhao, et al.
Published: (2025)
GatedLexiconNet: A Comprehensive End-to-End Handwritten Paragraph Text Recognition System
by: Kumari, Lalita, et al.
Published: (2024)
by: Kumari, Lalita, et al.
Published: (2024)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
by: Chen, Xinlei, et al.
Published: (2024)
by: Chen, Xinlei, et al.
Published: (2024)
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
by: Ma, Nanye, et al.
Published: (2024)
by: Ma, Nanye, et al.
Published: (2024)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Learning Camera Movement Control from Real-World Drone Videos
by: Hou, Yunzhong, et al.
Published: (2024)
by: Hou, Yunzhong, et al.
Published: (2024)
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
by: Zheng, Yupeng, et al.
Published: (2025)
by: Zheng, Yupeng, et al.
Published: (2025)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025)
by: Tang, Bingda, et al.
Published: (2025)
GaussianAD: Gaussian-Centric End-to-End Autonomous Driving
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
Codebook-Centric Deep Hashing: End-to-End Joint Learning of Semantic Hash Centers and Neural Hash Function
by: Yin, Shuo, et al.
Published: (2025)
by: Yin, Shuo, et al.
Published: (2025)
LDPM: Towards undersampled MRI reconstruction with MR-VAE and Latent Diffusion Prior
by: Tang, Xingjian, et al.
Published: (2024)
by: Tang, Xingjian, et al.
Published: (2024)
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
GIT-CXR: End-to-End Transformer for Chest X-Ray Report Generation
by: Sîrbu, Iustin, et al.
Published: (2025)
by: Sîrbu, Iustin, et al.
Published: (2025)
End-to-End HOI Reconstruction Transformer with Graph-based Encoding
by: Wang, Zhenrong, et al.
Published: (2025)
by: Wang, Zhenrong, et al.
Published: (2025)
End-to-End Fine-Tuning of 3D Texture Generation using Differentiable Rewards
by: Zamani, AmirHossein, et al.
Published: (2025)
by: Zamani, AmirHossein, et al.
Published: (2025)
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)
by: Li, Yingyan, et al.
Published: (2024)
Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-to-End Autonomous Driving
by: Xing, Zebin, et al.
Published: (2025)
by: Xing, Zebin, et al.
Published: (2025)
Probability Density Geodesics in Image Diffusion Latent Space
by: Yu, Qingtao, et al.
Published: (2025)
by: Yu, Qingtao, et al.
Published: (2025)
An End-to-End Depth-Based Pipeline for Selfie Image Rectification
by: Alhawwary, Ahmed, et al.
Published: (2024)
by: Alhawwary, Ahmed, et al.
Published: (2024)
Ultra-Efficient Decoding for End-to-End Neural Compression and Reconstruction
by: Rogers, Ethan G., et al.
Published: (2025)
by: Rogers, Ethan G., et al.
Published: (2025)
JLT: Clean-Latent Prediction in Latent Diffusion Transformers
by: Fu, Funing, et al.
Published: (2026)
by: Fu, Funing, et al.
Published: (2026)
Correcting Diffusion-Based Perceptual Image Compression with Privileged End-to-End Decoder
by: Ma, Yiyang, et al.
Published: (2024)
by: Ma, Yiyang, et al.
Published: (2024)
Transforming Hyperspectral Images Into Chemical Maps: A Novel End-to-End Deep Learning Approach
by: Engstrøm, Ole-Christian Galbo, et al.
Published: (2025)
by: Engstrøm, Ole-Christian Galbo, et al.
Published: (2025)
AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
by: Wang, Zile, et al.
Published: (2025)
by: Wang, Zile, et al.
Published: (2025)
Similar Items
-
What matters for Representation Alignment: Global Information or Spatial Structure?
by: Singh, Jaskirat, et al.
Published: (2025) -
SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
by: Zhao, Qinyu, et al.
Published: (2025) -
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025) -
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026) -
Improved Baselines with Representation Autoencoders
by: Singh, Jaskirat, et al.
Published: (2026)