Latent Diffusion Models with Masked AutoEncoders
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Junho, Shin, Jeongwoo, Choi, Hyungwook, Lee, Joonseok |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Guided Masked Autoencoder
by: Shin, Jeongwoo, et al.
Published: (2025)
by: Shin, Jeongwoo, et al.
Published: (2025)
Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schrödinger Bridge Matching
by: Shin, Jeongwoo, et al.
Published: (2026)
by: Shin, Jeongwoo, et al.
Published: (2026)
Scalable Frame Sampling for Video Classification: A Semi-Optimal Policy Approach with Reduced Search Space
by: Lee, Junho, et al.
Published: (2024)
by: Lee, Junho, et al.
Published: (2024)
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024)
by: Hahm, Jaehoon, et al.
Published: (2024)
Equivariant Latent Alignment via Flow Matching under Group Symmetries
by: Kim, Sunghyun, et al.
Published: (2026)
by: Kim, Sunghyun, et al.
Published: (2026)
Latent Enhancing AutoEncoder for Occluded Image Classification
by: Kotwal, Ketan, et al.
Published: (2024)
by: Kotwal, Ketan, et al.
Published: (2024)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
by: Labatie, Antoine, et al.
Published: (2025)
by: Labatie, Antoine, et al.
Published: (2025)
Distribution Matching Variational AutoEncoder
by: Ye, Sen, et al.
Published: (2025)
by: Ye, Sen, et al.
Published: (2025)
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026)
by: Lee, Junho, et al.
Published: (2026)
Robust Multimodal Survival Prediction with the Latent Differentiation Conditional Variational AutoEncoder
by: Zhou, Junjie, et al.
Published: (2025)
by: Zhou, Junjie, et al.
Published: (2025)
Occlusion-Aware 3D Hand-Object Pose Estimation with Masked AutoEncoders
by: Yang, Hui, et al.
Published: (2025)
by: Yang, Hui, et al.
Published: (2025)
Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark
by: Cao, Bing, et al.
Published: (2024)
by: Cao, Bing, et al.
Published: (2024)
3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud Pretraining
by: Yan, Siming, et al.
Published: (2023)
by: Yan, Siming, et al.
Published: (2023)
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning
by: Su, Kunming, et al.
Published: (2024)
by: Su, Kunming, et al.
Published: (2024)
Improved AutoEncoder with LSTM module and KL divergence
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
by: Lee, Inseo, et al.
Published: (2025)
by: Lee, Inseo, et al.
Published: (2025)
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026)
by: Lee, Hyun, et al.
Published: (2026)
ColorMAE: Exploring data-independent masking strategies in Masked AutoEncoders
by: Hinojosa, Carlos, et al.
Published: (2024)
by: Hinojosa, Carlos, et al.
Published: (2024)
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
by: Ha, Seongsu, et al.
Published: (2024)
by: Ha, Seongsu, et al.
Published: (2024)
Appearance Blur-driven AutoEncoder and Motion-guided Memory Module for Video Anomaly Detection
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
Detecting AI-Generated Images via Contextual Anomaly Estimation in Masked AutoEncoders
by: Jang, Minsuk, et al.
Published: (2025)
by: Jang, Minsuk, et al.
Published: (2025)
Detecting AutoEncoder is Enough to Catch LDM Generated Images
by: Vesnin, Dmitry, et al.
Published: (2024)
by: Vesnin, Dmitry, et al.
Published: (2024)
H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
SAEN-BGS: Energy-Efficient Spiking AutoEncoder Network for Background Subtraction
by: Zhang, Zhixuan, et al.
Published: (2025)
by: Zhang, Zhixuan, et al.
Published: (2025)
NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields
by: Irshad, Muhammad Zubair, et al.
Published: (2024)
by: Irshad, Muhammad Zubair, et al.
Published: (2024)
CE-VAE: Capsule Enhanced Variational AutoEncoder for Underwater Image Enhancement
by: Pucci, Rita, et al.
Published: (2024)
by: Pucci, Rita, et al.
Published: (2024)
AccuQuant: Simulating Multiple Denoising Steps for Quantizing Diffusion Models
by: Lee, Seunghoon, et al.
Published: (2025)
by: Lee, Seunghoon, et al.
Published: (2025)
Enhancing Creative Generation on Stable Diffusion-based Models
by: Han, Jiyeon, et al.
Published: (2025)
by: Han, Jiyeon, et al.
Published: (2025)
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
by: Kamenetsky, Ronen, et al.
Published: (2025)
by: Kamenetsky, Ronen, et al.
Published: (2025)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder
by: Bora, Maheswar, et al.
Published: (2024)
by: Bora, Maheswar, et al.
Published: (2024)
Physics Informed Capsule Enhanced Variational AutoEncoder for Underwater Image Enhancement
by: Martinel, Niki, et al.
Published: (2025)
by: Martinel, Niki, et al.
Published: (2025)
Towards Scalable Human-aligned Benchmark for Text-guided Image Editing
by: Ryu, Suho, et al.
Published: (2025)
by: Ryu, Suho, et al.
Published: (2025)
Modality-Aware Representation Learning for Zero-shot Sketch-based Image Retrieval
by: Lyou, Eunyi, et al.
Published: (2024)
by: Lyou, Eunyi, et al.
Published: (2024)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
Improving Diffusion Models for Authentic Virtual Try-on in the Wild
by: Choi, Yisol, et al.
Published: (2024)
by: Choi, Yisol, et al.
Published: (2024)
Multi-modal Intermediate Feature Interaction AutoEncoder for Overall Survival Prediction of Esophageal Squamous Cell Cancer
by: Wu, Chengyu, et al.
Published: (2024)
by: Wu, Chengyu, et al.
Published: (2024)
Similar Items
-
Self-Guided Masked Autoencoder
by: Shin, Jeongwoo, et al.
Published: (2025) -
Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schrödinger Bridge Matching
by: Shin, Jeongwoo, et al.
Published: (2026) -
Scalable Frame Sampling for Video Classification: A Semi-Optimal Policy Approach with Reduced Search Space
by: Lee, Junho, et al.
Published: (2024) -
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024) -
Equivariant Latent Alignment via Flow Matching under Group Symmetries
by: Kim, Sunghyun, et al.
Published: (2026)