VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Yifei, Wu, Xiaoshan, Hu, Xinting, Hu, Tao, Sun, Yangtian, Lyu, Xiaoyang, Wang, Bo, Ma, Lin, Ma, Yuewen, Wang, Zhongrui, Qi, Xiaojuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LiFR-Seg: Anytime High-Frame-Rate Segmentation via Event-Guided Propagation
von: Wu, Xiaoshan, et al.
Veröffentlicht: (2026)
von: Wu, Xiaoshan, et al.
Veröffentlicht: (2026)
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
EAG3R: Event-Augmented 3D Geometry Estimation for Dynamic and Extreme-Lighting Scenes
von: Wu, Xiaoshan, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoshan, et al.
Veröffentlicht: (2025)
Splatter a Video: Video Gaussian Representation for Versatile Processing
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2024)
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2024)
EX-4D: EXtreme Viewpoint 4D Video Synthesis via Depth Watertight Mesh
von: Hu, Tao, et al.
Veröffentlicht: (2025)
von: Hu, Tao, et al.
Veröffentlicht: (2025)
Stereo World Model: Camera-Guided Stereo Video Generation
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2026)
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2026)
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
Veda: Scalable Video Diffusion via Distilled Sparse Attention
von: Han, Shihao, et al.
Veröffentlicht: (2026)
von: Han, Shihao, et al.
Veröffentlicht: (2026)
VADMamba++: Efficient Video Anomaly Detection via Hybrid Modeling in Grayscale Space
von: Lyu, Jihao, et al.
Veröffentlicht: (2026)
von: Lyu, Jihao, et al.
Veröffentlicht: (2026)
VADMamba: Exploring State Space Models for Fast Video Anomaly Detection
von: Lyu, Jiahao, et al.
Veröffentlicht: (2025)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2025)
SeqTex: Generate Mesh Textures in Video Sequence
von: Yuan, Ze, et al.
Veröffentlicht: (2025)
von: Yuan, Ze, et al.
Veröffentlicht: (2025)
UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
von: Li, Peiming, et al.
Veröffentlicht: (2025)
von: Li, Peiming, et al.
Veröffentlicht: (2025)
TrackSSM: A General Motion Predictor by State-Space Model
von: Hu, Bin, et al.
Veröffentlicht: (2024)
von: Hu, Bin, et al.
Veröffentlicht: (2024)
RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
von: Zhu, Kai, et al.
Veröffentlicht: (2026)
von: Zhu, Kai, et al.
Veröffentlicht: (2026)
Spec-Gaussian: Anisotropic View-Dependent Appearance for 3D Gaussian Splatting
von: Yang, Ziyi, et al.
Veröffentlicht: (2024)
von: Yang, Ziyi, et al.
Veröffentlicht: (2024)
RoboSSM: Scalable In-context Imitation Learning via State-Space Models
von: Yoo, Youngju, et al.
Veröffentlicht: (2025)
von: Yoo, Youngju, et al.
Veröffentlicht: (2025)
VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos
von: Liu, Pengyiang, et al.
Veröffentlicht: (2026)
von: Liu, Pengyiang, et al.
Veröffentlicht: (2026)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
von: Liu, Kunhao, et al.
Veröffentlicht: (2025)
von: Liu, Kunhao, et al.
Veröffentlicht: (2025)
DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization
von: Ye, Zijian, et al.
Veröffentlicht: (2025)
von: Ye, Zijian, et al.
Veröffentlicht: (2025)
Speculative Decoding for Autoregressive Video Generation
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding
von: Li, Kangcong, et al.
Veröffentlicht: (2026)
von: Li, Kangcong, et al.
Veröffentlicht: (2026)
LIVE: Long-horizon Interactive Video World Modeling
von: Huang, Junchao, et al.
Veröffentlicht: (2026)
von: Huang, Junchao, et al.
Veröffentlicht: (2026)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
MAGI-1: Autoregressive Video Generation at Scale
von: ai, Sand., et al.
Veröffentlicht: (2025)
von: ai, Sand., et al.
Veröffentlicht: (2025)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
Time-SSM: Simplifying and Unifying State Space Models for Time Series Forecasting
von: Hu, Jiaxi, et al.
Veröffentlicht: (2024)
von: Hu, Jiaxi, et al.
Veröffentlicht: (2024)
Pathwise Test-Time Correction for Autoregressive Long Video Generation
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2026)
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2026)
DO3D: Self-supervised Learning of Decomposed Object-aware 3D Motion and Depth from Monocular Videos
von: Wu, Xiuzhe, et al.
Veröffentlicht: (2024)
von: Wu, Xiuzhe, et al.
Veröffentlicht: (2024)
LongSSM: On the Length Extension of State-space Models in Language Modelling
von: Wang, Shida
Veröffentlicht: (2024)
von: Wang, Shida
Veröffentlicht: (2024)
Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
von: Tang, Siao, et al.
Veröffentlicht: (2026)
von: Tang, Siao, et al.
Veröffentlicht: (2026)
VideoDirector: Precise Video Editing via Text-to-Video Models
von: Wang, Yukun, et al.
Veröffentlicht: (2024)
von: Wang, Yukun, et al.
Veröffentlicht: (2024)
Stateful Token Reduction for Long-Video Hybrid VLMs
von: Jiang, Jindong, et al.
Veröffentlicht: (2026)
von: Jiang, Jindong, et al.
Veröffentlicht: (2026)
AnchorSync: Global Consistency Optimization for Long Video Editing
von: Liu, Zichi, et al.
Veröffentlicht: (2025)
von: Liu, Zichi, et al.
Veröffentlicht: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
VFIMamba: Video Frame Interpolation with State Space Models
von: Zhang, Guozhen, et al.
Veröffentlicht: (2024)
von: Zhang, Guozhen, et al.
Veröffentlicht: (2024)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
Scaling RL to Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2025)
von: Chen, Yukang, et al.
Veröffentlicht: (2025)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
von: Ren, Shuhuai, et al.
Veröffentlicht: (2025)
von: Ren, Shuhuai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LiFR-Seg: Anytime High-Frame-Rate Segmentation via Event-Guided Propagation
von: Wu, Xiaoshan, et al.
Veröffentlicht: (2026) -
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
von: Wang, Bo, et al.
Veröffentlicht: (2025) -
EAG3R: Event-Augmented 3D Geometry Estimation for Dynamic and Extreme-Lighting Scenes
von: Wu, Xiaoshan, et al.
Veröffentlicht: (2025) -
Splatter a Video: Video Gaussian Representation for Versatile Processing
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2024) -
EX-4D: EXtreme Viewpoint 4D Video Synthesis via Depth Watertight Mesh
von: Hu, Tao, et al.
Veröffentlicht: (2025)