LongStream: Long-Sequence Streaming Autoregressive Visual Geometry
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Chong, Chen, Xianda, Xie, Tao, Yin, Wei, Ren, Weiqiang, Zhang, Qian, Guo, Xiaoyang, Wang, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
von: Cheng, Chong, et al.
Veröffentlicht: (2026)
von: Cheng, Chong, et al.
Veröffentlicht: (2026)
HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation
von: Zhang, Conglang, et al.
Veröffentlicht: (2026)
von: Zhang, Conglang, et al.
Veröffentlicht: (2026)
PAS3R: Pose-Adaptive Streaming 3D Reconstruction for Long Video Sequences
von: Xu, Lanbo, et al.
Veröffentlicht: (2026)
von: Xu, Lanbo, et al.
Veröffentlicht: (2026)
LONG3R: Long Sequence Streaming 3D Reconstruction
von: Chen, Zhuoguang, et al.
Veröffentlicht: (2025)
von: Chen, Zhuoguang, et al.
Veröffentlicht: (2025)
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
von: Hu, Yu, et al.
Veröffentlicht: (2025)
von: Hu, Yu, et al.
Veröffentlicht: (2025)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos
von: Liu, Pengyiang, et al.
Veröffentlicht: (2026)
von: Liu, Pengyiang, et al.
Veröffentlicht: (2026)
OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
von: Lu, Si-Yu, et al.
Veröffentlicht: (2026)
von: Lu, Si-Yu, et al.
Veröffentlicht: (2026)
StreamReady: Learning What to Answer and When in Long Streaming Videos
von: Azad, Shehreen, et al.
Veröffentlicht: (2026)
von: Azad, Shehreen, et al.
Veröffentlicht: (2026)
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026)
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026)
SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
von: Deng, Junyuan, et al.
Veröffentlicht: (2025)
von: Deng, Junyuan, et al.
Veröffentlicht: (2025)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
von: Zhen, Dingcheng, et al.
Veröffentlicht: (2025)
von: Zhen, Dingcheng, et al.
Veröffentlicht: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
von: Chen, Zhengqing, et al.
Veröffentlicht: (2025)
von: Chen, Zhengqing, et al.
Veröffentlicht: (2025)
Streaming 4D Visual Geometry Transformer
von: Zhuo, Dong, et al.
Veröffentlicht: (2025)
von: Zhuo, Dong, et al.
Veröffentlicht: (2025)
DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive Model
von: Chen, Bohong, et al.
Veröffentlicht: (2025)
von: Chen, Bohong, et al.
Veröffentlicht: (2025)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
von: Li, Ruibin, et al.
Veröffentlicht: (2026)
von: Li, Ruibin, et al.
Veröffentlicht: (2026)
MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
von: Zhao, Canyu, et al.
Veröffentlicht: (2024)
von: Zhao, Canyu, et al.
Veröffentlicht: (2024)
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
von: Xie, Tao, et al.
Veröffentlicht: (2026)
von: Xie, Tao, et al.
Veröffentlicht: (2026)
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
Entropy-Guided k-Guard Sampling for Long-Horizon Autoregressive Video Generation
von: Han, Yizhao, et al.
Veröffentlicht: (2026)
von: Han, Yizhao, et al.
Veröffentlicht: (2026)
Streaming Autoregressive Video Generation via Diagonal Distillation
von: Liu, Jinxiu, et al.
Veröffentlicht: (2026)
von: Liu, Jinxiu, et al.
Veröffentlicht: (2026)
FlowNar: Scalable Streaming Narration for Long-Form Videos
von: Zhong, Zeyun, et al.
Veröffentlicht: (2026)
von: Zhong, Zeyun, et al.
Veröffentlicht: (2026)
Boost 3D Reconstruction using Diffusion-based Monocular Camera Calibration
von: Deng, Junyuan, et al.
Veröffentlicht: (2024)
von: Deng, Junyuan, et al.
Veröffentlicht: (2024)
ViSpeak: Visual Instruction Feedback in Streaming Videos
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
von: Lyu, Hengye, et al.
Veröffentlicht: (2026)
von: Lyu, Hengye, et al.
Veröffentlicht: (2026)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2024)
von: Zhang, Haoji, et al.
Veröffentlicht: (2024)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
From Two-Stream to One-Stream: Efficient RGB-T Tracking via Mutual Prompt Learning and Knowledge Distillation
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
StreamingClaw Technical Report
von: Chen, Jiawei, et al.
Veröffentlicht: (2026)
von: Chen, Jiawei, et al.
Veröffentlicht: (2026)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
von: Jin, Xinqi, et al.
Veröffentlicht: (2025)
von: Jin, Xinqi, et al.
Veröffentlicht: (2025)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
von: Xie, Ming, et al.
Veröffentlicht: (2026)
von: Xie, Ming, et al.
Veröffentlicht: (2026)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation
von: Zhan, Yifan, et al.
Veröffentlicht: (2026)
von: Zhan, Yifan, et al.
Veröffentlicht: (2026)
MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology
von: Yang, Shu, et al.
Veröffentlicht: (2024)
von: Yang, Shu, et al.
Veröffentlicht: (2024)
Finding Visual Saliency in Continuous Spike Stream
von: Zhu, Lin, et al.
Veröffentlicht: (2024)
von: Zhu, Lin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
von: Cheng, Chong, et al.
Veröffentlicht: (2026) -
HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation
von: Zhang, Conglang, et al.
Veröffentlicht: (2026) -
PAS3R: Pose-Adaptive Streaming 3D Reconstruction for Long Video Sequences
von: Xu, Lanbo, et al.
Veröffentlicht: (2026) -
LONG3R: Long Sequence Streaming 3D Reconstruction
von: Chen, Zhuoguang, et al.
Veröffentlicht: (2025) -
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
von: Hu, Yu, et al.
Veröffentlicht: (2025)