Latte: Latent Diffusion Transformer for Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xin, Wang, Yaohui, Chen, Xinyuan, Jia, Gengyun, Liu, Ziwei, Li, Yuan-Fang, Chen, Cunjian, Qiao, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
LEO: Generative Latent Image Animator for Human Video Synthesis
by: Wang, Yaohui, et al.
Published: (2023)
by: Wang, Yaohui, et al.
Published: (2023)
Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
by: Ma, Xin, et al.
Published: (2025)
by: Ma, Xin, et al.
Published: (2025)
Training-free Stylized Text-to-Image Generation with Fast Inference
by: Ma, Xin, et al.
Published: (2025)
by: Ma, Xin, et al.
Published: (2025)
4Diffusion: Multi-view Video Diffusion Model for 4D Generation
by: Zhang, Haiyu, et al.
Published: (2024)
by: Zhang, Haiyu, et al.
Published: (2024)
CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
by: Wu, Xiaoxue, et al.
Published: (2025)
by: Wu, Xiaoxue, et al.
Published: (2025)
AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
by: Zhang, Haiyu, et al.
Published: (2025)
by: Zhang, Haiyu, et al.
Published: (2025)
ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
by: Wu, Xiaoxue, et al.
Published: (2025)
by: Wu, Xiaoxue, et al.
Published: (2025)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
by: Peng, Bo, et al.
Published: (2023)
by: Peng, Bo, et al.
Published: (2023)
LIA-X: Interpretable Latent Portrait Animator
by: Wang, Yaohui, et al.
Published: (2025)
by: Wang, Yaohui, et al.
Published: (2025)
PARE: Pruning and Adaptive Routing for Efficient Video Generation
by: Wang, Yutong, et al.
Published: (2026)
by: Wang, Yutong, et al.
Published: (2026)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
by: Wang, Yutong, et al.
Published: (2025)
by: Wang, Yutong, et al.
Published: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion
by: Wen, Hao, et al.
Published: (2024)
by: Wen, Hao, et al.
Published: (2024)
RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
by: Huang, Ziqi, et al.
Published: (2024)
by: Huang, Ziqi, et al.
Published: (2024)
Vlogger: Make Your Dream A Vlog
by: Zhuang, Shaobin, et al.
Published: (2024)
by: Zhuang, Shaobin, et al.
Published: (2024)
Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
by: Fan, Weichen, et al.
Published: (2025)
by: Fan, Weichen, et al.
Published: (2025)
AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
by: Xu, Longhuan, et al.
Published: (2025)
by: Xu, Longhuan, et al.
Published: (2025)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
by: Li, Chunyu, et al.
Published: (2024)
by: Li, Chunyu, et al.
Published: (2024)
StructLDM: Structured Latent Diffusion for 3D Human Generation
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
VGDFR: Diffusion-based Video Generation with Dynamic Latent Frame Rate
by: Yuan, Zhihang, et al.
Published: (2025)
by: Yuan, Zhihang, et al.
Published: (2025)
MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation
by: Sun, Mingzhen, et al.
Published: (2024)
by: Sun, Mingzhen, et al.
Published: (2024)
Controlling the Latent Diffusion Model for Generative Image Shadow Removal via Residual Generation
by: Li, Xinjie, et al.
Published: (2024)
by: Li, Xinjie, et al.
Published: (2024)
TALL: Thumbnail Layout for Deepfake Video Detection
by: Xu, Yuting, et al.
Published: (2023)
by: Xu, Yuting, et al.
Published: (2023)
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
by: Li, Shujia, et al.
Published: (2025)
by: Li, Shujia, et al.
Published: (2025)
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025)
by: Si, Chenyang, et al.
Published: (2025)
EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained Diffusion
by: Huang, Zehuan, et al.
Published: (2023)
by: Huang, Zehuan, et al.
Published: (2023)
Video Generation with Predictive Latents
by: Zhao, Yian, et al.
Published: (2026)
by: Zhao, Yian, et al.
Published: (2026)
Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial Frame
by: Cao, Qinglong, et al.
Published: (2024)
by: Cao, Qinglong, et al.
Published: (2024)
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
by: Hu, Yaosi, et al.
Published: (2023)
by: Hu, Yaosi, et al.
Published: (2023)
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
by: Gao, Jianxiong, et al.
Published: (2025)
by: Gao, Jianxiong, et al.
Published: (2025)
CPA: Camera-pose-awareness Diffusion Transformer for Video Generation
by: Wang, Yuelei, et al.
Published: (2024)
by: Wang, Yuelei, et al.
Published: (2024)
Causal Debiasing for Visual Commonsense Reasoning
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
by: Yu, Ruonan, et al.
Published: (2026)
by: Yu, Ruonan, et al.
Published: (2026)
MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
by: Zhang, Yuechen, et al.
Published: (2025)
by: Zhang, Yuechen, et al.
Published: (2025)
TimeStep Master: Asymmetrical Mixture of Timestep LoRA Experts for Versatile and Efficient Diffusion Models in Vision
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion
by: Liu, Xian, et al.
Published: (2023)
by: Liu, Xian, et al.
Published: (2023)
Similar Items
-
Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models
by: Ma, Xin, et al.
Published: (2024) -
LEO: Generative Latent Image Animator for Human Video Synthesis
by: Wang, Yaohui, et al.
Published: (2023) -
Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
by: Ma, Xin, et al.
Published: (2025) -
Training-free Stylized Text-to-Image Generation with Fast Inference
by: Ma, Xin, et al.
Published: (2025) -
4Diffusion: Multi-view Video Diffusion Model for 4D Generation
by: Zhang, Haiyu, et al.
Published: (2024)