StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Tianrui, Li, Zhi, Yang, Shuo, Xi, Haocheng, Li, Muyang, Li, Xiuyu, Zhang, Lvmin, Yang, Keting, Peng, Kelly, Han, Song, Agrawala, Maneesh, Keutzer, Kurt, Kodaira, Akio, Xu, Chenfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
Transparent Image Layer Diffusion using Latent Transparency
by: Zhang, Lvmin, et al.
Published: (2024)
by: Zhang, Lvmin, et al.
Published: (2024)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
View-oriented Conversation Compiler for Agent Trace Analysis
by: Zhang, Lvmin, et al.
Published: (2026)
by: Zhang, Lvmin, et al.
Published: (2026)
Immiscible Diffusion: Accelerating Diffusion Training with Noise Assignment
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
StreamDiffusion: A Pipeline-level Solution for Real-time Interactive Generation
by: Kodaira, Akio, et al.
Published: (2023)
by: Kodaira, Akio, et al.
Published: (2023)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
Pretraining Frame Preservation for Lightweight Autoregressive Video History Embedding
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
Taming Flow-based I2V Models for Creative Video Editing
by: Kong, Xianghao, et al.
Published: (2025)
by: Kong, Xianghao, et al.
Published: (2025)
StreamDiT: Real-Time Streaming Text-to-Video Generation
by: Kodaira, Akio, et al.
Published: (2025)
by: Kodaira, Akio, et al.
Published: (2025)
Instance Segmentation of Scene Sketches Using Natural Image Priors
by: Tang, Mia, et al.
Published: (2025)
by: Tang, Mia, et al.
Published: (2025)
Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility
by: Li, Yiheng, et al.
Published: (2025)
by: Li, Yiheng, et al.
Published: (2025)
MoVer: Motion Verification for Motion Graphics Animations
by: Ma, Jiaju, et al.
Published: (2025)
by: Ma, Jiaju, et al.
Published: (2025)
Magic-Me: Identity-Specific Video Customized Diffusion
by: Ma, Ze, et al.
Published: (2024)
by: Ma, Ze, et al.
Published: (2024)
CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
Mixture of Contexts for Long Video Generation
by: Cai, Shengqu, et al.
Published: (2025)
by: Cai, Shengqu, et al.
Published: (2025)
HallE-Control: Controlling Object Hallucination in Large Multimodal Models
by: Zhai, Bohan, et al.
Published: (2023)
by: Zhai, Bohan, et al.
Published: (2023)
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
by: Yang, Yang, et al.
Published: (2026)
by: Yang, Yang, et al.
Published: (2026)
A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision
by: Peng, Chensheng, et al.
Published: (2024)
by: Peng, Chensheng, et al.
Published: (2024)
Captain Cinema: Towards Short Movie Generation
by: Xiao, Junfei, et al.
Published: (2025)
by: Xiao, Junfei, et al.
Published: (2025)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Mode Seeking meets Mean Seeking for Fast Long Video Generation
by: Cai, Shengqu, et al.
Published: (2026)
by: Cai, Shengqu, et al.
Published: (2026)
SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing
by: Zhou, Xuanyi, et al.
Published: (2026)
by: Zhou, Xuanyi, et al.
Published: (2026)
Segment Any Motion in Videos
by: Huang, Nan, et al.
Published: (2025)
by: Huang, Nan, et al.
Published: (2025)
StreamChat: Chatting with Streaming Video
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Adaptive 3D Gaussian Splatting Video Streaming
by: Gong, Han, et al.
Published: (2025)
by: Gong, Han, et al.
Published: (2025)
ScriptViz: A Visualization Tool to Aid Scriptwriting based on a Large Movie Database
by: Rao, Anyi, et al.
Published: (2024)
by: Rao, Anyi, et al.
Published: (2024)
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
by: Khaki, Samir, et al.
Published: (2025)
by: Khaki, Samir, et al.
Published: (2025)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
by: Lin, Junming, et al.
Published: (2024)
by: Lin, Junming, et al.
Published: (2024)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
by: Li, Jie, et al.
Published: (2023)
by: Li, Jie, et al.
Published: (2023)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
by: Jiang, Jindong, et al.
Published: (2025)
by: Jiang, Jindong, et al.
Published: (2025)
DeformStream: Deformation-based Adaptive Volumetric Video Streaming
by: Li, Boyan, et al.
Published: (2024)
by: Li, Boyan, et al.
Published: (2024)
Residual Context Diffusion Language Models
by: Hu, Yuezhou, et al.
Published: (2026)
by: Hu, Yuezhou, et al.
Published: (2026)
Dynamic Control Analysis of Various Side‐Stream Quaternary Extractive Distillation Configurations
by: Min Li, et al.
Published: (2024)
by: Min Li, et al.
Published: (2024)
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
by: Luo, Yawen, et al.
Published: (2026)
by: Luo, Yawen, et al.
Published: (2026)
Similar Items
-
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
by: Liang, Feng, et al.
Published: (2024) -
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
by: Zhang, Lvmin, et al.
Published: (2025) -
Transparent Image Layer Diffusion using Latent Transparency
by: Zhang, Lvmin, et al.
Published: (2024) -
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
by: Li, Xingyang, et al.
Published: (2025) -
View-oriented Conversation Compiler for Agent Trace Analysis
by: Zhang, Lvmin, et al.
Published: (2026)