Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Tianhao, Yuan, Jianlong, Feng, Wanquan, Fang, Shancheng, Liu, Jiawei, Zhou, SiYu, He, Qian, Xie, Hongtao, Zhang, Yongdong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations
by: Qi, Tianhao, et al.
Published: (2024)
by: Qi, Tianhao, et al.
Published: (2024)
AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion
by: Sun, Mingzhen, et al.
Published: (2025)
by: Sun, Mingzhen, et al.
Published: (2025)
GenMask: Adapting DiT for Segmentation via Direct Mask Generation
by: Yang, Yuhuan, et al.
Published: (2026)
by: Yang, Yuhuan, et al.
Published: (2026)
DiT4Edit: Diffusion Transformer for Image Editing
by: Feng, Kunyu, et al.
Published: (2024)
by: Feng, Kunyu, et al.
Published: (2024)
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
by: Fang, Jiarui, et al.
Published: (2024)
by: Fang, Jiarui, et al.
Published: (2024)
HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
by: Liu, Wenxuan, et al.
Published: (2024)
by: Liu, Wenxuan, et al.
Published: (2024)
FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation
by: Fang, Xueji, et al.
Published: (2026)
by: Fang, Xueji, et al.
Published: (2026)
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
by: Sun, Xibo, et al.
Published: (2024)
by: Sun, Xibo, et al.
Published: (2024)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
by: Chen, Sixiang, et al.
Published: (2025)
by: Chen, Sixiang, et al.
Published: (2025)
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
by: Lyu, Hengye, et al.
Published: (2026)
by: Lyu, Hengye, et al.
Published: (2026)
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
by: Yang, Lianwei, et al.
Published: (2025)
by: Yang, Lianwei, et al.
Published: (2025)
LaVin-DiT: Large Vision Diffusion Transformer
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
by: Ma, Teli, et al.
Published: (2026)
by: Ma, Teli, et al.
Published: (2026)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
IGD: Instructional Graphic Design with Multimodal Layer Generation
by: Qu, Yadong, et al.
Published: (2025)
by: Qu, Yadong, et al.
Published: (2025)
Diffusion-Guided Mask-Consistent Paired Mixing for Endoscopic Image Segmentation
by: Jie, Pengyu, et al.
Published: (2025)
by: Jie, Pengyu, et al.
Published: (2025)
TQ-DiT: Efficient Time-Aware Quantization for Diffusion Transformers
by: Hwang, Younghye, et al.
Published: (2025)
by: Hwang, Younghye, et al.
Published: (2025)
PTQ4DiT: Post-training Quantization for Diffusion Transformers
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
TaQ-DiT: Time-aware Quantization for Diffusion Transformers
by: Liu, Xinyan, et al.
Published: (2024)
by: Liu, Xinyan, et al.
Published: (2024)
DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression
by: Shi, Junqi, et al.
Published: (2026)
by: Shi, Junqi, et al.
Published: (2026)
DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution
by: Duan, Zheng-Peng, et al.
Published: (2025)
by: Duan, Zheng-Peng, et al.
Published: (2025)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation
by: Krishnamurthy, Bharath, et al.
Published: (2026)
by: Krishnamurthy, Bharath, et al.
Published: (2026)
VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers
by: Zheng, Jun, et al.
Published: (2024)
by: Zheng, Jun, et al.
Published: (2024)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
by: Guo, Junrong, et al.
Published: (2026)
by: Guo, Junrong, et al.
Published: (2026)
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
by: Xie, Yu, et al.
Published: (2025)
by: Xie, Yu, et al.
Published: (2025)
TextFlux: An OCR‐Free DiT Model for High‐Fidelity Multilingual Scene Text Synthesis
by: Yu Xie, et al.
Published: (2026)
by: Yu Xie, et al.
Published: (2026)
U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation
by: Wu, Linzhi, et al.
Published: (2025)
by: Wu, Linzhi, et al.
Published: (2025)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
by: Zhu, Rui, et al.
Published: (2024)
by: Zhu, Rui, et al.
Published: (2024)
DiT-JSCC: Rethinking Deep JSCC with Diffusion Transformers and Semantic Representations
by: Tan, Kailin, et al.
Published: (2026)
by: Tan, Kailin, et al.
Published: (2026)
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
DiVE: DiT-based Video Generation with Enhanced Control
by: Jiang, Junpeng, et al.
Published: (2024)
by: Jiang, Junpeng, et al.
Published: (2024)
I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
by: Feng, Wanquan, et al.
Published: (2024)
by: Feng, Wanquan, et al.
Published: (2024)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
by: Sun, Jiahui, et al.
Published: (2025)
by: Sun, Jiahui, et al.
Published: (2025)
Similar Items
-
DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations
by: Qi, Tianhao, et al.
Published: (2024) -
AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion
by: Sun, Mingzhen, et al.
Published: (2025) -
GenMask: Adapting DiT for Segmentation via Direct Mask Generation
by: Yang, Yuhuan, et al.
Published: (2026) -
DiT4Edit: Diffusion Transformer for Image Editing
by: Feng, Kunyu, et al.
Published: (2024) -
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
by: Fang, Gongfan, et al.
Published: (2024)