Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xiaoce, Zhou, Sifan, Wang, Kaifei, Xu, Leli, Qiu, Xuerui, He, Tao, Li, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
von: Cai, Xin, et al.
Veröffentlicht: (2026)
von: Cai, Xin, et al.
Veröffentlicht: (2026)
FD-DiT: Frequency Domain-Directed Diffusion Transformer for Low-Dose CT Reconstruction
von: Liu, Qiqing, et al.
Veröffentlicht: (2025)
von: Liu, Qiqing, et al.
Veröffentlicht: (2025)
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
von: Ma, Teli, et al.
Veröffentlicht: (2026)
von: Ma, Teli, et al.
Veröffentlicht: (2026)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
von: Mikaeili, Aryan, et al.
Veröffentlicht: (2026)
von: Mikaeili, Aryan, et al.
Veröffentlicht: (2026)
DiVE: DiT-based Video Generation with Enhanced Control
von: Jiang, Junpeng, et al.
Veröffentlicht: (2024)
von: Jiang, Junpeng, et al.
Veröffentlicht: (2024)
Slit-Induced Reflection Mode Conversion Between Fundamental Lamb Modes in Elastic Plates with Low-Frequency and Broadband Response
von: Feng, Kaifei, et al.
Veröffentlicht: (2026)
von: Feng, Kaifei, et al.
Veröffentlicht: (2026)
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
DiT4Edit: Diffusion Transformer for Image Editing
von: Feng, Kunyu, et al.
Veröffentlicht: (2024)
von: Feng, Kunyu, et al.
Veröffentlicht: (2024)
LaVin-DiT: Large Vision Diffusion Transformer
von: Wang, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Wang, Zhaoqing, et al.
Veröffentlicht: (2024)
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
von: Zhou, Chao, et al.
Veröffentlicht: (2026)
von: Zhou, Chao, et al.
Veröffentlicht: (2026)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
von: Lin, Hangyu, et al.
Veröffentlicht: (2026)
von: Lin, Hangyu, et al.
Veröffentlicht: (2026)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
von: Fang, Jiarui, et al.
Veröffentlicht: (2024)
DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression
von: Shi, Junqi, et al.
Veröffentlicht: (2026)
von: Shi, Junqi, et al.
Veröffentlicht: (2026)
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
Decoupled Alignment for Robust Plug-and-Play Adaptation
von: Luo, Haozheng, et al.
Veröffentlicht: (2024)
von: Luo, Haozheng, et al.
Veröffentlicht: (2024)
PromptLoop: Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
von: Lee, Suhyeon, et al.
Veröffentlicht: (2025)
von: Lee, Suhyeon, et al.
Veröffentlicht: (2025)
TaQ-DiT: Time-aware Quantization for Diffusion Transformers
von: Liu, Xinyan, et al.
Veröffentlicht: (2024)
von: Liu, Xinyan, et al.
Veröffentlicht: (2024)
Density-Informed VAE (DiVAE): Reliable Log-Prior Probability via Density Alignment Regularization
von: Alessi, Michele, et al.
Veröffentlicht: (2025)
von: Alessi, Michele, et al.
Veröffentlicht: (2025)
HiMat: DiT-based Ultra-High Resolution SVBRDF Generation
von: Wang, Zixiong, et al.
Veröffentlicht: (2025)
von: Wang, Zixiong, et al.
Veröffentlicht: (2025)
PTQ4DiT: Post-training Quantization for Diffusion Transformers
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal
von: Wu, Chenyang, et al.
Veröffentlicht: (2026)
von: Wu, Chenyang, et al.
Veröffentlicht: (2026)
MaterialPicker: Multi-Modal DiT-Based Material Generation
von: Ma, Xiaohe, et al.
Veröffentlicht: (2024)
von: Ma, Xiaohe, et al.
Veröffentlicht: (2024)
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
von: Huo, Yanru, et al.
Veröffentlicht: (2025)
von: Huo, Yanru, et al.
Veröffentlicht: (2025)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation
von: Wu, Linzhi, et al.
Veröffentlicht: (2025)
von: Wu, Linzhi, et al.
Veröffentlicht: (2025)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
von: Niu, Zhikang, et al.
Veröffentlicht: (2025)
von: Niu, Zhikang, et al.
Veröffentlicht: (2025)
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
von: Tian, Yuchuan, et al.
Veröffentlicht: (2024)
von: Tian, Yuchuan, et al.
Veröffentlicht: (2024)
Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones
von: Chen, Ranfei, et al.
Veröffentlicht: (2025)
von: Chen, Ranfei, et al.
Veröffentlicht: (2025)
Ortho-Hydra: Orthogonalized Experts for DiT LoRA
von: Ji, Seunghyun
Veröffentlicht: (2026)
von: Ji, Seunghyun
Veröffentlicht: (2026)
DiT-JSCC: Rethinking Deep JSCC with Diffusion Transformers and Semantic Representations
von: Tan, Kailin, et al.
Veröffentlicht: (2026)
von: Tan, Kailin, et al.
Veröffentlicht: (2026)
GenMask: Adapting DiT for Segmentation via Direct Mask Generation
von: Yang, Yuhuan, et al.
Veröffentlicht: (2026)
von: Yang, Yuhuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
von: Cai, Xin, et al.
Veröffentlicht: (2026) -
FD-DiT: Frequency Domain-Directed Diffusion Transformer for Low-Dose CT Reconstruction
von: Liu, Qiqing, et al.
Veröffentlicht: (2025) -
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
von: Ma, Teli, et al.
Veröffentlicht: (2026) -
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026) -
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)