Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ruibin, Yang, Tao, Ai, Fangzhou, Wu, Tianhe, Wen, Shilei, Peng, Bingyue, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
by: Li, Ruibin, et al.
Published: (2025)
by: Li, Ruibin, et al.
Published: (2025)
Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
by: Wu, Tianhe, et al.
Published: (2026)
by: Wu, Tianhe, et al.
Published: (2026)
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
by: Cheng, Chong, et al.
Published: (2026)
by: Cheng, Chong, et al.
Published: (2026)
Streaming Autoregressive Video Generation via Diagonal Distillation
by: Liu, Jinxiu, et al.
Published: (2026)
by: Liu, Jinxiu, et al.
Published: (2026)
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
by: Guo, Yanjun, et al.
Published: (2026)
by: Guo, Yanjun, et al.
Published: (2026)
Multi-Level Decoupled Relational Distillation for Heterogeneous Architectures
by: Yang, Yaoxin, et al.
Published: (2025)
by: Yang, Yaoxin, et al.
Published: (2025)
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation
by: Liu, Yuheng, et al.
Published: (2026)
by: Liu, Yuheng, et al.
Published: (2026)
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
by: Qu, Jinyuan, et al.
Published: (2024)
by: Qu, Jinyuan, et al.
Published: (2024)
CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
by: Yan, Shilin, et al.
Published: (2025)
by: Yan, Shilin, et al.
Published: (2025)
Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality Assumption
by: Chen, Du, et al.
Published: (2025)
by: Chen, Du, et al.
Published: (2025)
STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation
by: Wang, Jiamin, et al.
Published: (2025)
by: Wang, Jiamin, et al.
Published: (2025)
Veda: Scalable Video Diffusion via Distilled Sparse Attention
by: Han, Shihao, et al.
Published: (2026)
by: Han, Shihao, et al.
Published: (2026)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
Towards One-step Causal Video Generation via Adversarial Self-Distillation
by: Yang, Yongqi, et al.
Published: (2025)
by: Yang, Yongqi, et al.
Published: (2025)
Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
by: Bendel, Matthew, et al.
Published: (2026)
by: Bendel, Matthew, et al.
Published: (2026)
Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer
by: Liu, Wenxuan, et al.
Published: (2025)
by: Liu, Wenxuan, et al.
Published: (2025)
Train Short, Inference Long: Training-free Horizon Extension for Autoregressive Video Generation
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Waver: Wave Your Way to Lifelike Video Generation
by: Zhang, Yifu, et al.
Published: (2025)
by: Zhang, Yifu, et al.
Published: (2025)
RORem: Training a Robust Object Remover with Human-in-the-Loop
by: Li, Ruibin, et al.
Published: (2025)
by: Li, Ruibin, et al.
Published: (2025)
Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
by: Lu, Yunhong, et al.
Published: (2025)
by: Lu, Yunhong, et al.
Published: (2025)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
Generative Refinement Networks for Visual Synthesis
by: Han, Jian, et al.
Published: (2026)
by: Han, Jian, et al.
Published: (2026)
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
by: Lu, Yu, et al.
Published: (2024)
by: Lu, Yu, et al.
Published: (2024)
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
MedHorizon: Towards Long-context Medical Video Understanding in the Wild
by: Du, Bodong, et al.
Published: (2026)
by: Du, Bodong, et al.
Published: (2026)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank
by: Wu, Tianhe, et al.
Published: (2025)
by: Wu, Tianhe, et al.
Published: (2025)
Dual-Stream Spectral Decoupling Distillation for Remote Sensing Object Detection
by: Gao, Xiangyi, et al.
Published: (2025)
by: Gao, Xiangyi, et al.
Published: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
by: Wang, Junxi, et al.
Published: (2026)
by: Wang, Junxi, et al.
Published: (2026)
Entropy-Guided k-Guard Sampling for Long-Horizon Autoregressive Video Generation
by: Han, Yizhao, et al.
Published: (2026)
by: Han, Yizhao, et al.
Published: (2026)
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
by: Wang, Boyuan, et al.
Published: (2025)
by: Wang, Boyuan, et al.
Published: (2025)
StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation
by: Jiao, Guanlong, et al.
Published: (2026)
by: Jiao, Guanlong, et al.
Published: (2026)
Similar Items
-
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
by: Li, Ruibin, et al.
Published: (2025) -
Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
by: Wu, Tianhe, et al.
Published: (2026) -
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
by: Cheng, Chong, et al.
Published: (2026) -
Streaming Autoregressive Video Generation via Diagonal Distillation
by: Liu, Jinxiu, et al.
Published: (2026) -
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
by: Guo, Yanjun, et al.
Published: (2026)