Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wuyang, Pan, Wentao, Luan, Po-Chien, Gao, Yang, Alahi, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models
by: Luan, Po-Chien, et al.
Published: (2026)
by: Luan, Po-Chien, et al.
Published: (2026)
Drift-Resistant Navigation World Model with Anchored Epipolar Guidance
by: Luan, Po-Chien, et al.
Published: (2026)
by: Luan, Po-Chien, et al.
Published: (2026)
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
by: Li, Wuyang, et al.
Published: (2026)
by: Li, Wuyang, et al.
Published: (2026)
Multi-Transmotion: Pre-trained Model for Human Motion Prediction
by: Gao, Yang, et al.
Published: (2024)
by: Gao, Yang, et al.
Published: (2024)
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
by: Hassan, Mariam, et al.
Published: (2025)
by: Hassan, Mariam, et al.
Published: (2025)
Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
by: Hassan, Mariam, et al.
Published: (2026)
by: Hassan, Mariam, et al.
Published: (2026)
Unified Human Localization and Trajectory Prediction with Monocular Vision
by: Luan, Po-Chien, et al.
Published: (2025)
by: Luan, Po-Chien, et al.
Published: (2025)
OmniTraj: Pre-Training on Heterogeneous Data for Adaptive and Zero-Shot Human Trajectory Prediction
by: Gao, Yang, et al.
Published: (2025)
by: Gao, Yang, et al.
Published: (2025)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
VoxDet: Rethinking 3D Semantic Occupancy Prediction as Dense Object Detection
by: Li, Wuyang, et al.
Published: (2025)
by: Li, Wuyang, et al.
Published: (2025)
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
by: Yesiltepe, Hidir, et al.
Published: (2025)
by: Yesiltepe, Hidir, et al.
Published: (2025)
Sim-to-Real Causal Transfer: A Metric Learning Approach to Causally-Aware Interaction Representations
by: Rahimi, Ahmad, et al.
Published: (2023)
by: Rahimi, Ahmad, et al.
Published: (2023)
SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
by: Wang, Jianyi, et al.
Published: (2025)
by: Wang, Jianyi, et al.
Published: (2025)
Social-Pose: Enhancing Trajectory Prediction with Human Body Pose
by: Gao, Yang, et al.
Published: (2025)
by: Gao, Yang, et al.
Published: (2025)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
by: Yang, Shaoshu, et al.
Published: (2025)
by: Yang, Shaoshu, et al.
Published: (2025)
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
by: Elmoghany, Mohamed, et al.
Published: (2026)
by: Elmoghany, Mohamed, et al.
Published: (2026)
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
by: Yi, Hongwei, et al.
Published: (2025)
by: Yi, Hongwei, et al.
Published: (2025)
RAP: 3D Rasterization Augmented End-to-End Planning
by: Feng, Lan, et al.
Published: (2025)
by: Feng, Lan, et al.
Published: (2025)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
by: Yang, Ying, et al.
Published: (2026)
by: Yang, Ying, et al.
Published: (2026)
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
Infinite Gaze Generation for Videos with Autoregressive Diffusion
by: Kang, Jenna, et al.
Published: (2026)
by: Kang, Jenna, et al.
Published: (2026)
FG$^2$: Fine-Grained Cross-View Localization by Fine-Grained Feature Matching
by: Xia, Zimin, et al.
Published: (2025)
by: Xia, Zimin, et al.
Published: (2025)
Endora: Video Generation Models as Endoscopy Simulators
by: Li, Chenxin, et al.
Published: (2024)
by: Li, Chenxin, et al.
Published: (2024)
Social-Transmotion: Promptable Human Trajectory Prediction
by: Saadatnejad, Saeed, et al.
Published: (2023)
by: Saadatnejad, Saeed, et al.
Published: (2023)
NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
by: Nakka, Krishna Kanth, et al.
Published: (2025)
by: Nakka, Krishna Kanth, et al.
Published: (2025)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
Representation Recycling for Streaming Video Analysis
by: Ertenli, Can Ufuk, et al.
Published: (2022)
by: Ertenli, Can Ufuk, et al.
Published: (2022)
RAIN: Real-time Animation of Infinite Video Stream
by: Shu, Zhilei, et al.
Published: (2024)
by: Shu, Zhilei, et al.
Published: (2024)
X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
by: Feng, X., et al.
Published: (2026)
by: Feng, X., et al.
Published: (2026)
From Ideal to Real: Stable Video Object Removal under Imperfect Conditions
by: Hu, Jiagao, et al.
Published: (2026)
by: Hu, Jiagao, et al.
Published: (2026)
SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
by: Haghighi, Yasaman, et al.
Published: (2026)
by: Haghighi, Yasaman, et al.
Published: (2026)
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
by: Messaoud, Kaouther, et al.
Published: (2025)
by: Messaoud, Kaouther, et al.
Published: (2025)
Loc$^2$: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching
by: Xia, Zimin, et al.
Published: (2025)
by: Xia, Zimin, et al.
Published: (2025)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026)
by: Xiong, Tianwei, et al.
Published: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
Controllable Generative Video Compression
by: Ding, Ding, et al.
Published: (2026)
by: Ding, Ding, et al.
Published: (2026)
VideoTetris: Towards Compositional Text-to-Video Generation
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
Similar Items
-
Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models
by: Luan, Po-Chien, et al.
Published: (2026) -
Drift-Resistant Navigation World Model with Anchored Epipolar Guidance
by: Luan, Po-Chien, et al.
Published: (2026) -
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
by: Li, Wuyang, et al.
Published: (2026) -
Multi-Transmotion: Pre-trained Model for Human Motion Prediction
by: Gao, Yang, et al.
Published: (2024) -
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
by: Hassan, Mariam, et al.
Published: (2025)