DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Xin, You, Zhiyuan, Zhang, Zhoutong, Xue, Tianfan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
AutoDIR: Automatic All-in-One Image Restoration with Latent Diffusion
by: Jiang, Yitong, et al.
Published: (2023)
by: Jiang, Yitong, et al.
Published: (2023)
Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
by: Zhu, Lunjie, et al.
Published: (2026)
by: Zhu, Lunjie, et al.
Published: (2026)
Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset
by: You, Zhiyuan, et al.
Published: (2024)
by: You, Zhiyuan, et al.
Published: (2024)
Learning to Refocus with Video Diffusion Models
by: Tedla, SaiKiran, et al.
Published: (2025)
by: Tedla, SaiKiran, et al.
Published: (2025)
UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
PhoCoLens: Photorealistic and Consistent Reconstruction in Lensless Imaging
by: Cai, Xin, et al.
Published: (2024)
by: Cai, Xin, et al.
Published: (2024)
Eliminating VAE for Fast and High-Resolution Generative Detail Restoration
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
PhotoAgent: Agentic Photo Editing with Exploratory Visual Aesthetic Planning
by: Yao, Mingde, et al.
Published: (2026)
by: Yao, Mingde, et al.
Published: (2026)
Plug-and-play Diffusion Models for Image Compressive Sensing with Data Consistency Projection
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
Revisiting the Generalization Problem of Low-level Vision Models Through the Lens of Image Deraining
by: Hu, Jinfan, et al.
Published: (2025)
by: Hu, Jinfan, et al.
Published: (2025)
Depicting Beyond Scores: Advancing Image Quality Assessment through Multi-modal Language Models
by: You, Zhiyuan, et al.
Published: (2023)
by: You, Zhiyuan, et al.
Published: (2023)
DiffIER: Optimizing Diffusion Models with Iterative Error Reduction
by: Chen, Ao, et al.
Published: (2025)
by: Chen, Ao, et al.
Published: (2025)
Detail-Preserving Latent Diffusion for Stable Shadow Removal
by: Xu, Jiamin, et al.
Published: (2024)
by: Xu, Jiamin, et al.
Published: (2024)
FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
by: Zhuang, Junhao, et al.
Published: (2025)
by: Zhuang, Junhao, et al.
Published: (2025)
PromptLoop: Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
by: Lee, Suhyeon, et al.
Published: (2025)
by: Lee, Suhyeon, et al.
Published: (2025)
Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging
by: Cai, Chang, et al.
Published: (2025)
by: Cai, Chang, et al.
Published: (2025)
LoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning
by: Gao, Chenjian, et al.
Published: (2025)
by: Gao, Chenjian, et al.
Published: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
Improved Video VAE for Latent Video Diffusion Model
by: Wu, Pingyu, et al.
Published: (2024)
by: Wu, Pingyu, et al.
Published: (2024)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
by: Yuan, Zhihang, et al.
Published: (2025)
by: Yuan, Zhihang, et al.
Published: (2025)
LenslessFace: An End-to-End Optimized Lensless System for Privacy-Preserving Face Verification
by: Cai, Xin, et al.
Published: (2024)
by: Cai, Xin, et al.
Published: (2024)
Boosting Latent Diffusion Models via Disentangled Representation Alignment
by: Page, John, et al.
Published: (2026)
by: Page, John, et al.
Published: (2026)
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
by: Bi, Tianci, et al.
Published: (2025)
by: Bi, Tianci, et al.
Published: (2025)
Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
by: Kui, Xiaoyan, et al.
Published: (2025)
by: Kui, Xiaoyan, et al.
Published: (2025)
The Devil is in the Details: Boosting Guided Depth Super-Resolution via Rethinking Cross-Modal Alignment and Aggregation
by: Jiang, Xinni, et al.
Published: (2024)
by: Jiang, Xinni, et al.
Published: (2024)
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
by: Leng, Xingjian, et al.
Published: (2025)
by: Leng, Xingjian, et al.
Published: (2025)
LDPM: Towards undersampled MRI reconstruction with MR-VAE and Latent Diffusion Prior
by: Tang, Xingjian, et al.
Published: (2024)
by: Tang, Xingjian, et al.
Published: (2024)
Latent Dirichlet Transformer VAE for Hyperspectral Unmixing with Bundled Endmembers
by: Giannetti, Giancarlo, et al.
Published: (2025)
by: Giannetti, Giancarlo, et al.
Published: (2025)
FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
by: FSVideo Team, et al.
Published: (2026)
by: FSVideo Team, et al.
Published: (2026)
MTC-VAE: Multi-Level Temporal Compression with Content Awareness
by: Dong, Yubo, et al.
Published: (2026)
by: Dong, Yubo, et al.
Published: (2026)
Detail Matters: Mamba-Inspired Joint Unfolding Network for Snapshot Spectral Compressive Imaging
by: Qin, Mengjie, et al.
Published: (2025)
by: Qin, Mengjie, et al.
Published: (2025)
LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter
by: Li, Runyi, et al.
Published: (2025)
by: Li, Runyi, et al.
Published: (2025)
Towards Extreme Image Compression with Latent Feature Guidance and Diffusion Prior
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
Plug-and-Play Versatile Compressed Video Enhancement
by: Zeng, Huimin, et al.
Published: (2025)
by: Zeng, Huimin, et al.
Published: (2025)
AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
by: Chen, Yutian, et al.
Published: (2026)
by: Chen, Yutian, et al.
Published: (2026)
Similar Items
-
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025) -
Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution
by: You, Zhiyuan, et al.
Published: (2025) -
AutoDIR: Automatic All-in-One Image Restoration with Latent Diffusion
by: Jiang, Yitong, et al.
Published: (2023) -
Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
by: Wang, Xiaoce, et al.
Published: (2026) -
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
by: Zhu, Lunjie, et al.
Published: (2026)