Diffusion Autoencoders are Scalable Image Tokenizers
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yinbo, Girdhar, Rohit, Wang, Xiaolong, Rambhatla, Sai Saketh, Misra, Ishan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
SelfEval: Leveraging the discriminative nature of generative models for evaluation
by: Rambhatla, Sai Saketh, et al.
Published: (2023)
by: Rambhatla, Sai Saketh, et al.
Published: (2023)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023)
by: Girdhar, Rohit, et al.
Published: (2023)
Generating Illustrated Instructions
by: Menon, Sachit, et al.
Published: (2023)
by: Menon, Sachit, et al.
Published: (2023)
LLMs can see and hear without any training
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
MotiF: Making Text Count in Image Animation with Motion Focal Loss
by: Wang, Shijie, et al.
Published: (2024)
by: Wang, Shijie, et al.
Published: (2024)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Consistent Flow Distillation for Text-to-3D Generation
by: Yan, Runjie, et al.
Published: (2025)
by: Yan, Runjie, et al.
Published: (2025)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
by: Singh, Mannat, et al.
Published: (2023)
by: Singh, Mannat, et al.
Published: (2023)
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Improving the Diffusability of Autoencoders
by: Skorokhodov, Ivan, et al.
Published: (2025)
by: Skorokhodov, Ivan, et al.
Published: (2025)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024)
by: Surkov, Viacheslav, et al.
Published: (2024)
Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology
by: Kraus, Oren, et al.
Published: (2024)
by: Kraus, Oren, et al.
Published: (2024)
TD-MPC2: Scalable, Robust World Models for Continuous Control
by: Hansen, Nicklas, et al.
Published: (2023)
by: Hansen, Nicklas, et al.
Published: (2023)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
by: Yeung, Calvin, et al.
Published: (2026)
by: Yeung, Calvin, et al.
Published: (2026)
Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers
by: Crowson, Katherine, et al.
Published: (2024)
by: Crowson, Katherine, et al.
Published: (2024)
Hyperspherical Autoencoder for High-Fidelity Image Reconstruction and Generation
by: Chang, Hun, et al.
Published: (2026)
by: Chang, Hun, et al.
Published: (2026)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
by: Ghasemi, Narges, et al.
Published: (2025)
by: Ghasemi, Narges, et al.
Published: (2025)
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
Language-Guided Image Tokenization for Generation
by: Zha, Kaiwen, et al.
Published: (2024)
by: Zha, Kaiwen, et al.
Published: (2024)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Accelerating Diffusion Transformers with Token-wise Feature Caching
by: Zou, Chang, et al.
Published: (2024)
by: Zou, Chang, et al.
Published: (2024)
Wavelet-Driven Generalizable Framework for Deepfake Face Forgery Detection
by: Baru, Lalith Bharadwaj, et al.
Published: (2024)
by: Baru, Lalith Bharadwaj, et al.
Published: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
Communication-Inspired Tokenization for Structured Image Representations
by: Davtyan, Aram, et al.
Published: (2026)
by: Davtyan, Aram, et al.
Published: (2026)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026)
by: Lee, Hyun, et al.
Published: (2026)
DMin: Scalable Training Data Influence Estimation for Diffusion Models
by: Lin, Huawei, et al.
Published: (2024)
by: Lin, Huawei, et al.
Published: (2024)
SceneTok: A Compressed, Diffusable Token Space for 3D Scenes
by: Asim, Mohammad, et al.
Published: (2026)
by: Asim, Mohammad, et al.
Published: (2026)
Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
Single-pass Adaptive Image Tokenization for Minimum Program Search
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
Diffuse and Disperse: Image Generation with Representation Regularization
by: Wang, Runqian, et al.
Published: (2025)
by: Wang, Runqian, et al.
Published: (2025)
Identifying Bias in Deep Neural Networks Using Image Transforms
by: Erukude, Sai Teja, et al.
Published: (2024)
by: Erukude, Sai Teja, et al.
Published: (2024)
ST-GDance++: A Scalable Spatial-Temporal Diffusion for Long-Duration Group Choreography
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
by: Zou, Chang, et al.
Published: (2024)
by: Zou, Chang, et al.
Published: (2024)
ToDo: Token Downsampling for Efficient Generation of High-Resolution Images
by: Smith, Ethan, et al.
Published: (2024)
by: Smith, Ethan, et al.
Published: (2024)
Trajectory-aligned Space-time Tokens for Few-shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2024)
by: Kumar, Pulkit, et al.
Published: (2024)
Similar Items
-
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024) -
SelfEval: Leveraging the discriminative nature of generative models for evaluation
by: Rambhatla, Sai Saketh, et al.
Published: (2023) -
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023) -
Generating Illustrated Instructions
by: Menon, Sachit, et al.
Published: (2023) -
LLMs can see and hear without any training
by: Ashutosh, Kumar, et al.
Published: (2025)