UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Min, Zhu, Hongzhou, Wang, Yingze, Yan, Bokai, Zhang, Jintao, He, Guande, Yang, Ling, Li, Chongxuan, Zhu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
by: Zhao, Min, et al.
Published: (2024)
by: Zhao, Min, et al.
Published: (2024)
Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
by: Bao, Fan, et al.
Published: (2024)
by: Bao, Fan, et al.
Published: (2024)
Consistency Diffusion Bridge Models
by: He, Guande, et al.
Published: (2024)
by: He, Guande, et al.
Published: (2024)
ViViD: Video Virtual Try-on using Diffusion Models
by: Fang, Zixun, et al.
Published: (2024)
by: Fang, Zixun, et al.
Published: (2024)
Novel View Extrapolation with Video Diffusion Priors
by: Liu, Kunhao, et al.
Published: (2024)
by: Liu, Kunhao, et al.
Published: (2024)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
by: Huang, Xun, et al.
Published: (2025)
by: Huang, Xun, et al.
Published: (2025)
Causality in Video Diffusers is Separable from Denoising
by: Bai, Xingjian, et al.
Published: (2026)
by: Bai, Xingjian, et al.
Published: (2026)
6Bit-Diffusion: Inference-Time Mixed-Precision Quantization for Video Diffusion Models
by: Su, Rundong, et al.
Published: (2026)
by: Su, Rundong, et al.
Published: (2026)
Elucidating the Preconditioning in Consistency Distillation
by: Zheng, Kaiwen, et al.
Published: (2025)
by: Zheng, Kaiwen, et al.
Published: (2025)
Scaling Diffusion Transformers Efficiently via $μ$P
by: Zheng, Chenyu, et al.
Published: (2025)
by: Zheng, Chenyu, et al.
Published: (2025)
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
by: Yu, Ruonan, et al.
Published: (2026)
by: Yu, Ruonan, et al.
Published: (2026)
DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
by: Lu, Cheng, et al.
Published: (2022)
by: Lu, Cheng, et al.
Published: (2022)
Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models
by: Cui, Peng, et al.
Published: (2024)
by: Cui, Peng, et al.
Published: (2024)
ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
by: Wu, Yunfeng, et al.
Published: (2026)
by: Wu, Yunfeng, et al.
Published: (2026)
TRecViT: A Recurrent Video Transformer
by: Pătrăucean, Viorica, et al.
Published: (2024)
by: Pătrăucean, Viorica, et al.
Published: (2024)
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
by: Zhao, Chengfeng, et al.
Published: (2026)
by: Zhao, Chengfeng, et al.
Published: (2026)
PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
by: Du, Shian, et al.
Published: (2025)
by: Du, Shian, et al.
Published: (2025)
DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
by: Issachar, Noam, et al.
Published: (2025)
by: Issachar, Noam, et al.
Published: (2025)
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
ViTMAlis: Towards Latency-Critical Mobile Video Analytics with Vision Transformers
by: Zhang, Miao, et al.
Published: (2026)
by: Zhang, Miao, et al.
Published: (2026)
LoopViT: Scaling Visual ARC with Looped Transformers
by: Shu, Wen-Jie, et al.
Published: (2026)
by: Shu, Wen-Jie, et al.
Published: (2026)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
Pose-Aware Diffusion for 3D Generation
by: Zhou, Zihan, et al.
Published: (2026)
by: Zhou, Zihan, et al.
Published: (2026)
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
by: Zhu, Wentao
Published: (2024)
by: Zhu, Wentao
Published: (2024)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control
by: Zhong, Yong, et al.
Published: (2024)
by: Zhong, Yong, et al.
Published: (2024)
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
by: Lu, Beijia, et al.
Published: (2025)
by: Lu, Beijia, et al.
Published: (2025)
ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models
by: Gao, Kaifeng, et al.
Published: (2024)
by: Gao, Kaifeng, et al.
Published: (2024)
MoViE: Mobile Diffusion for Video Editing
by: Karjauv, Adil, et al.
Published: (2024)
by: Karjauv, Adil, et al.
Published: (2024)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
by: Rajabi, Javad, et al.
Published: (2026)
by: Rajabi, Javad, et al.
Published: (2026)
ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
LoViT: Long Video Transformer for Surgical Phase Recognition
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
Improved Video VAE for Latent Video Diffusion Model
by: Wu, Pingyu, et al.
Published: (2024)
by: Wu, Pingyu, et al.
Published: (2024)
Similar Items
-
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025) -
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025) -
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
by: Zhao, Min, et al.
Published: (2026) -
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
by: Zhao, Min, et al.
Published: (2026) -
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
by: Zhao, Min, et al.
Published: (2024)