Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Tianhao, Yuan, Jianlong, Feng, Wanquan, Fang, Shancheng, Liu, Jiawei, Zhou, SiYu, He, Qian, Xie, Hongtao, Zhang, Yongdong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916662231433216
author Qi, Tianhao
Yuan, Jianlong
Feng, Wanquan
Fang, Shancheng
Liu, Jiawei
Zhou, SiYu
He, Qian
Xie, Hongtao
Zhang, Yongdong
author_facet Qi, Tianhao
Yuan, Jianlong
Feng, Wanquan
Fang, Shancheng
Liu, Jiawei
Zhou, SiYu
He, Qian
Xie, Hongtao
Zhang, Yongdong
contents Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask$^2$DiT, a novel approach that establishes fine-grained, one-to-one alignment between video segments and their corresponding text annotations. Specifically, we introduce a symmetric binary mask at each attention layer within the DiT architecture, ensuring that each text annotation applies exclusively to its respective video segment while preserving temporal coherence across visual tokens. This attention mechanism enables precise segment-level textual-to-visual alignment, allowing the DiT architecture to effectively handle video generation tasks with a fixed number of scenes. To further equip the DiT architecture with the ability to generate additional scenes based on existing ones, we incorporate a segment-level conditional mask, which conditions each newly generated segment on the preceding video segments, thereby enabling auto-regressive scene extension. Both qualitative and quantitative experiments confirm that Mask$^2$DiT excels in maintaining visual consistency across segments while ensuring semantic alignment between each segment and its corresponding text description. Our project page is https://tianhao-qi.github.io/Mask2DiTProject.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19881
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
Qi, Tianhao
Yuan, Jianlong
Feng, Wanquan
Fang, Shancheng
Liu, Jiawei
Zhou, SiYu
He, Qian
Xie, Hongtao
Zhang, Yongdong
Computer Vision and Pattern Recognition
Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask$^2$DiT, a novel approach that establishes fine-grained, one-to-one alignment between video segments and their corresponding text annotations. Specifically, we introduce a symmetric binary mask at each attention layer within the DiT architecture, ensuring that each text annotation applies exclusively to its respective video segment while preserving temporal coherence across visual tokens. This attention mechanism enables precise segment-level textual-to-visual alignment, allowing the DiT architecture to effectively handle video generation tasks with a fixed number of scenes. To further equip the DiT architecture with the ability to generate additional scenes based on existing ones, we incorporate a segment-level conditional mask, which conditions each newly generated segment on the preceding video segments, thereby enabling auto-regressive scene extension. Both qualitative and quantitative experiments confirm that Mask$^2$DiT excels in maintaining visual consistency across segments while ensuring semantic alignment between each segment and its corresponding text description. Our project page is https://tianhao-qi.github.io/Mask2DiTProject.
title Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.19881