Decoupling Complexity from Scale in Latent Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Tianxiong, Tian, Xingye, Wang, Xuebo, Jiang, Boyuan, Tao, Xin, Wan, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
by: Zhong, Tianxiong, et al.
Published: (2025)
by: Zhong, Tianxiong, et al.
Published: (2025)
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
by: Hu, Jiahao, et al.
Published: (2024)
by: Hu, Jiahao, et al.
Published: (2024)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
by: Cheng, Junhao, et al.
Published: (2026)
by: Cheng, Junhao, et al.
Published: (2026)
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
by: Ge, Wenhang, et al.
Published: (2026)
by: Ge, Wenhang, et al.
Published: (2026)
DiffHarmony: Latent Diffusion Model Meets Image Harmonization
by: Zhou, Pengfei, et al.
Published: (2024)
by: Zhou, Pengfei, et al.
Published: (2024)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution
by: Li, Hanting, et al.
Published: (2025)
by: Li, Hanting, et al.
Published: (2025)
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
Latent Diffusion Model without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
FLDM-VTON: Faithful Latent Diffusion Model for Virtual Try-on
by: Wang, Chenhui, et al.
Published: (2024)
by: Wang, Chenhui, et al.
Published: (2024)
RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image Generation
by: Cao, Boyuan, et al.
Published: (2024)
by: Cao, Boyuan, et al.
Published: (2024)
Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer
by: Li, Boyuan, et al.
Published: (2024)
by: Li, Boyuan, et al.
Published: (2024)
Hybrid Latent Reasoning with Decoupled Policy Optimization
by: Cheng, Tao, et al.
Published: (2026)
by: Cheng, Tao, et al.
Published: (2026)
Small Object Detection in Complex Backgrounds with Multi-Scale Attention and Global Relation Modeling
by: Tao, Wenguang, et al.
Published: (2026)
by: Tao, Wenguang, et al.
Published: (2026)
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
by: Wang, Yuan, et al.
Published: (2026)
by: Wang, Yuan, et al.
Published: (2026)
DDT: Decoupled Diffusion Transformer
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
by: Hou, Liang, et al.
Published: (2025)
by: Hou, Liang, et al.
Published: (2025)
Anatomically Guided Latent Diffusion for Brain MRI Progression Modeling
by: Wan, Cheng, et al.
Published: (2026)
by: Wan, Cheng, et al.
Published: (2026)
MSF: Efficient Diffusion Model Via Multi-Scale Latent Factorize
by: Xu, Haohang, et al.
Published: (2025)
by: Xu, Haohang, et al.
Published: (2025)
Efficient Video Diffusion Models: Advancements and Challenges
by: Shao, Shitong, et al.
Published: (2026)
by: Shao, Shitong, et al.
Published: (2026)
Terra: Explorable Native 3D World Model with Point Latents
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion
by: Yue, Zhengrong, et al.
Published: (2026)
by: Yue, Zhengrong, et al.
Published: (2026)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
Nested Diffusion Models Using Hierarchical Latent Priors
by: Zhang, Xiao, et al.
Published: (2024)
by: Zhang, Xiao, et al.
Published: (2024)
Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
by: Luo, Liwei, et al.
Published: (2025)
by: Luo, Liwei, et al.
Published: (2025)
FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing
by: Jiang, Yilei, et al.
Published: (2025)
by: Jiang, Yilei, et al.
Published: (2025)
PIDNet: Progressive Implicit Decouple Network for Multimodal Action Quality Assessment
by: Li, Qiqi, et al.
Published: (2026)
by: Li, Qiqi, et al.
Published: (2026)
Unpaired Deblurring via Decoupled Diffusion Model
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
A-SDM: Accelerating Stable Diffusion through Model Assembly and Feature Inheritance Strategies
by: Zhu, Jinchao, et al.
Published: (2024)
by: Zhu, Jinchao, et al.
Published: (2024)
Diffusion As Self-Distillation: End-to-End Latent Diffusion In One Model
by: Wang, Xiyuan, et al.
Published: (2025)
by: Wang, Xiyuan, et al.
Published: (2025)
Learning Latent Space Hierarchical EBM Diffusion Models
by: Cui, Jiali, et al.
Published: (2024)
by: Cui, Jiali, et al.
Published: (2024)
LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models
by: Jiang, Hai, et al.
Published: (2024)
by: Jiang, Hai, et al.
Published: (2024)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Models
by: Bradbury, Rowan, et al.
Published: (2025)
by: Bradbury, Rowan, et al.
Published: (2025)
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
by: Xian, Jia Jun Cheng, et al.
Published: (2025)
by: Xian, Jia Jun Cheng, et al.
Published: (2025)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
by: Yang, Zhenhao, et al.
Published: (2026)
by: Yang, Zhenhao, et al.
Published: (2026)
LatentHDR: Decoupling Exposure from Diffusion via Conditional Latent-to-Latent Mapping for Text/Image-to-Panoramic HDR
by: Fekri, Pedram, et al.
Published: (2026)
by: Fekri, Pedram, et al.
Published: (2026)
DeCo: Decoupled Human-Centered Diffusion Video Editing with Motion Consistency
by: Zhong, Xiaojing, et al.
Published: (2024)
by: Zhong, Xiaojing, et al.
Published: (2024)
Similar Items
-
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
by: Zhong, Tianxiong, et al.
Published: (2025) -
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
by: Hu, Jiahao, et al.
Published: (2024) -
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
by: Cheng, Junhao, et al.
Published: (2026) -
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
by: Xiang, Xunzhi, et al.
Published: (2025) -
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
by: Ge, Wenhang, et al.
Published: (2026)