BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zeyu, Chang, Shuning, He, Yuanyu, Han, Yizeng, Tang, Jiasheng, Wang, Fan, Zhuang, Bohan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
por: Inferix Team, et al.
Publicado: (2025)
por: Inferix Team, et al.
Publicado: (2025)
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
por: Liu, Akide, et al.
Publicado: (2025)
por: Liu, Akide, et al.
Publicado: (2025)
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
por: Pu, Yifan, et al.
Publicado: (2025)
por: Pu, Yifan, et al.
Publicado: (2025)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
por: Chen, Zhuokun, et al.
Publicado: (2026)
por: Chen, Zhuokun, et al.
Publicado: (2026)
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
por: Zhao, Wangbo, et al.
Publicado: (2025)
por: Zhao, Wangbo, et al.
Publicado: (2025)
SparseDiT: Token Sparsification for Efficient Diffusion Transformer
por: Chang, Shuning, et al.
Publicado: (2024)
por: Chang, Shuning, et al.
Publicado: (2024)
LongVLM: Efficient Long Video Understanding via Large Language Models
por: Weng, Yuetian, et al.
Publicado: (2024)
por: Weng, Yuetian, et al.
Publicado: (2024)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
por: Xiao, Xinyu, et al.
Publicado: (2026)
por: Xiao, Xinyu, et al.
Publicado: (2026)
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
por: Gu, Youping, et al.
Publicado: (2025)
por: Gu, Youping, et al.
Publicado: (2025)
Dynamic Diffusion Transformer
por: Zhao, Wangbo, et al.
Publicado: (2024)
por: Zhao, Wangbo, et al.
Publicado: (2024)
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
por: Zhao, Wangbo, et al.
Publicado: (2025)
por: Zhao, Wangbo, et al.
Publicado: (2025)
BachVid: Training-Free Video Generation with Consistent Background and Character
por: Yan, Han, et al.
Publicado: (2025)
por: Yan, Han, et al.
Publicado: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
por: Qiu, Zongyang, et al.
Publicado: (2025)
por: Qiu, Zongyang, et al.
Publicado: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
por: Fang, Ye, et al.
Publicado: (2025)
por: Fang, Ye, et al.
Publicado: (2025)
DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
por: Di, Donglin, et al.
Publicado: (2024)
por: Di, Donglin, et al.
Publicado: (2024)
Motion Mamba: Efficient and Long Sequence Motion Generation
por: Zhang, Zeyu, et al.
Publicado: (2024)
por: Zhang, Zeyu, et al.
Publicado: (2024)
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
por: Chen, Junsong, et al.
Publicado: (2025)
por: Chen, Junsong, et al.
Publicado: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
por: Qin, Bosheng, et al.
Publicado: (2023)
por: Qin, Bosheng, et al.
Publicado: (2023)
UniVid: The Open-Source Unified Video Model
por: Luo, Jiabin, et al.
Publicado: (2025)
por: Luo, Jiabin, et al.
Publicado: (2025)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
por: Chen, Feng, et al.
Publicado: (2025)
por: Chen, Feng, et al.
Publicado: (2025)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
por: Chen, Feng, et al.
Publicado: (2024)
por: Chen, Feng, et al.
Publicado: (2024)
ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality
por: He, Yefei, et al.
Publicado: (2024)
por: He, Yefei, et al.
Publicado: (2024)
VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors
por: Tang, Jimin, et al.
Publicado: (2026)
por: Tang, Jimin, et al.
Publicado: (2026)
Neighboring Autoregressive Modeling for Efficient Visual Generation
por: He, Yefei, et al.
Publicado: (2025)
por: He, Yefei, et al.
Publicado: (2025)
Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
por: Cui, Justin, et al.
Publicado: (2025)
por: Cui, Justin, et al.
Publicado: (2025)
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
por: Zhao, Wangbo, et al.
Publicado: (2024)
por: Zhao, Wangbo, et al.
Publicado: (2024)
MLV-Edit: Towards Consistent and Highly Efficient Editing for Minute-Level Videos
por: Cao, Yangyi, et al.
Publicado: (2026)
por: Cao, Yangyi, et al.
Publicado: (2026)
Minute-Long Videos with Dual Parallelisms
por: Wang, Zeqing, et al.
Publicado: (2025)
por: Wang, Zeqing, et al.
Publicado: (2025)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
por: Yang, Zhenhao, et al.
Publicado: (2026)
por: Yang, Zhenhao, et al.
Publicado: (2026)
BWCache: Accelerating Video Diffusion Transformers through Block-Wise Caching
por: Cui, Hanshuai, et al.
Publicado: (2025)
por: Cui, Hanshuai, et al.
Publicado: (2025)
DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
por: Jin, Wonjoon, et al.
Publicado: (2026)
por: Jin, Wonjoon, et al.
Publicado: (2026)
MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion
por: Susladkar, Onkar, et al.
Publicado: (2024)
por: Susladkar, Onkar, et al.
Publicado: (2024)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
por: Liang, Feng, et al.
Publicado: (2023)
por: Liang, Feng, et al.
Publicado: (2023)
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
por: Li, Hui, et al.
Publicado: (2024)
por: Li, Hui, et al.
Publicado: (2024)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
por: Wu, Jianzong, et al.
Publicado: (2025)
por: Wu, Jianzong, et al.
Publicado: (2025)
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
por: Zhang, Zeyu, et al.
Publicado: (2024)
por: Zhang, Zeyu, et al.
Publicado: (2024)
AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling
por: Mai, Ziyang, et al.
Publicado: (2026)
por: Mai, Ziyang, et al.
Publicado: (2026)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
por: Yuan, Hangjie, et al.
Publicado: (2025)
por: Yuan, Hangjie, et al.
Publicado: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
por: Wang, Junke, et al.
Publicado: (2024)
por: Wang, Junke, et al.
Publicado: (2024)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
por: Wang, Yuqing, et al.
Publicado: (2024)
por: Wang, Yuqing, et al.
Publicado: (2024)
Ejemplares similares
-
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
por: Inferix Team, et al.
Publicado: (2025) -
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
por: Liu, Akide, et al.
Publicado: (2025) -
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
por: Pu, Yifan, et al.
Publicado: (2025) -
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
por: Chen, Zhuokun, et al.
Publicado: (2026) -
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
por: Zhao, Wangbo, et al.
Publicado: (2025)