LayerT2V: A Unified Multi-Layer Video Generation Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Guangzhao, Cen, Kangrui, Zhao, Baixuan, Xin, Yi, Luo, Siqi, Zhai, Guangtao, Zhang, Lei, Liu, Xiaohong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915819351441408
author Li, Guangzhao
Cen, Kangrui
Zhao, Baixuan
Xin, Yi
Luo, Siqi
Zhai, Guangtao
Zhang, Lei
Liu, Xiaohong
author_facet Li, Guangzhao
Cen, Kangrui
Zhao, Baixuan
Xin, Yi
Luo, Siqi
Zhai, Guangtao
Zhang, Lei
Liu, Xiaohong
contents Text-to-video generation has advanced rapidly, but existing methods typically output only the final composited video and lack editable layered representations, limiting their use in professional workflows. We propose \textbf{LayerT2V}, a unified multi-layer video generation framework that produces multiple semantically consistent outputs in a single inference pass: the full video, an independent background layer, and multiple foreground RGB layers with corresponding alpha mattes. Our key insight is that recent video generation backbones use high compression in both time and space, enabling us to serialize multiple layer representations along the temporal dimension and jointly model them on a shared generation trajectory. This turns cross-layer consistency into an intrinsic objective, improving semantic alignment and temporal coherence. To mitigate layer ambiguity and conditional leakage, we augment a shared DiT backbone with LayerAdaLN and layer-aware cross-attention modulation. LayerT2V is trained in three stages: alpha mask VAE adaptation, joint multi-layer learning, and multi-foreground extension. We also introduce \textbf{VidLayer}, the first large-scale dataset for multi-layer video generation. Extensive experiments demonstrate that LayerT2V substantially outperforms prior methods in visual fidelity, temporal consistency, and cross-layer coherence.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04228
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LayerT2V: A Unified Multi-Layer Video Generation Framework
Li, Guangzhao
Cen, Kangrui
Zhao, Baixuan
Xin, Yi
Luo, Siqi
Zhai, Guangtao
Zhang, Lei
Liu, Xiaohong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Text-to-video generation has advanced rapidly, but existing methods typically output only the final composited video and lack editable layered representations, limiting their use in professional workflows. We propose \textbf{LayerT2V}, a unified multi-layer video generation framework that produces multiple semantically consistent outputs in a single inference pass: the full video, an independent background layer, and multiple foreground RGB layers with corresponding alpha mattes. Our key insight is that recent video generation backbones use high compression in both time and space, enabling us to serialize multiple layer representations along the temporal dimension and jointly model them on a shared generation trajectory. This turns cross-layer consistency into an intrinsic objective, improving semantic alignment and temporal coherence. To mitigate layer ambiguity and conditional leakage, we augment a shared DiT backbone with LayerAdaLN and layer-aware cross-attention modulation. LayerT2V is trained in three stages: alpha mask VAE adaptation, joint multi-layer learning, and multi-foreground extension. We also introduce \textbf{VidLayer}, the first large-scale dataset for multi-layer video generation. Extensive experiments demonstrate that LayerT2V substantially outperforms prior methods in visual fidelity, temporal consistency, and cross-layer coherence.
title LayerT2V: A Unified Multi-Layer Video Generation Framework
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
url https://arxiv.org/abs/2508.04228