Diffusion Transformers with Representation Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Boyang, Ma, Nanye, Tong, Shengbang, Xie, Saining |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
Flow Map Distillation Without Data
by: Tong, Shangyuan, et al.
Published: (2025)
by: Tong, Shangyuan, et al.
Published: (2025)
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
by: Ma, Nanye, et al.
Published: (2024)
by: Ma, Nanye, et al.
Published: (2024)
Improved Baselines with Representation Autoencoders
by: Singh, Jaskirat, et al.
Published: (2026)
by: Singh, Jaskirat, et al.
Published: (2026)
Transition Matching Distillation for Fast Video Generation
by: Nie, Weili, et al.
Published: (2026)
by: Nie, Weili, et al.
Published: (2026)
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
by: Leng, Xingjian, et al.
Published: (2025)
by: Leng, Xingjian, et al.
Published: (2025)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
by: Chen, Xinlei, et al.
Published: (2024)
by: Chen, Xinlei, et al.
Published: (2024)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025)
by: Tang, Bingda, et al.
Published: (2025)
MeanFlow Transformers with Representation Autoencoders
by: Hu, Zheyuan, et al.
Published: (2025)
by: Hu, Zheyuan, et al.
Published: (2025)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
by: Zhai, Yuexiang, et al.
Published: (2024)
by: Zhai, Yuexiang, et al.
Published: (2024)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
by: Chu, Tianzhe, et al.
Published: (2023)
by: Chu, Tianzhe, et al.
Published: (2023)
Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling
by: Xie, Tianyu, et al.
Published: (2025)
by: Xie, Tianyu, et al.
Published: (2025)
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023)
by: Yu, Yaodong, et al.
Published: (2023)
Robust Data Clustering with Outliers via Transformed Tensor Low-Rank Representation
by: Wu, Tong
Published: (2023)
by: Wu, Tong
Published: (2023)
Robust Representation Learning in Masked Autoencoders
by: Shrivastava, Anika, et al.
Published: (2026)
by: Shrivastava, Anika, et al.
Published: (2026)
What matters for Representation Alignment: Global Information or Spatial Structure?
by: Singh, Jaskirat, et al.
Published: (2025)
by: Singh, Jaskirat, et al.
Published: (2025)
Canonical Latent Representations in Conditional Diffusion Models
by: Xu, Yitao, et al.
Published: (2025)
by: Xu, Yitao, et al.
Published: (2025)
SARMAE: Masked Autoencoder for SAR Representation Learning
by: Liu, Danxu, et al.
Published: (2025)
by: Liu, Danxu, et al.
Published: (2025)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
by: Ma, Nanye, et al.
Published: (2025)
by: Ma, Nanye, et al.
Published: (2025)
Probing the Representational Power of Sparse Autoencoders in Vision Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
Sparse Autoencoders for Interpretable Medical Image Representation Learning
by: Wesp, Philipp, et al.
Published: (2026)
by: Wesp, Philipp, et al.
Published: (2026)
Attention-Guided Masked Autoencoders For Learning Image Representations
by: Sick, Leon, et al.
Published: (2024)
by: Sick, Leon, et al.
Published: (2024)
Mixed Autoencoder for Self-supervised Visual Representation Learning
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
by: Kumar, Amandeep, et al.
Published: (2026)
by: Kumar, Amandeep, et al.
Published: (2026)
Self-Refining Video Sampling
by: Jang, Sangwon, et al.
Published: (2026)
by: Jang, Sangwon, et al.
Published: (2026)
Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications
by: Roßteutscher, Immanuel, et al.
Published: (2025)
by: Roßteutscher, Immanuel, et al.
Published: (2025)
Improving the Diffusability of Autoencoders
by: Skorokhodov, Ivan, et al.
Published: (2025)
by: Skorokhodov, Ivan, et al.
Published: (2025)
CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models
by: He, Zhenghao, et al.
Published: (2026)
by: He, Zhenghao, et al.
Published: (2026)
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
Counterfactual Explanations for Medical Image Classification and Regression using Diffusion Autoencoder
by: Atad, Matan, et al.
Published: (2024)
by: Atad, Matan, et al.
Published: (2024)
Mode Seeking meets Mean Seeking for Fast Long Video Generation
by: Cai, Shengqu, et al.
Published: (2026)
by: Cai, Shengqu, et al.
Published: (2026)
Token Caching for Diffusion Transformer Acceleration
by: Lou, Jinming, et al.
Published: (2024)
by: Lou, Jinming, et al.
Published: (2024)
MUST: Modality-Specific Representation-Aware Transformer for Diffusion-Enhanced Survival Prediction with Missing Modality
by: Kim, Kyungwon, et al.
Published: (2026)
by: Kim, Kyungwon, et al.
Published: (2026)
MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Diffusing Differentiable Representations
by: Savani, Yash, et al.
Published: (2024)
by: Savani, Yash, et al.
Published: (2024)
Similar Items
-
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026) -
Flow Map Distillation Without Data
by: Tong, Shangyuan, et al.
Published: (2025) -
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
by: Ma, Nanye, et al.
Published: (2024) -
Improved Baselines with Representation Autoencoders
by: Singh, Jaskirat, et al.
Published: (2026) -
Transition Matching Distillation for Fast Video Generation
by: Nie, Weili, et al.
Published: (2026)