Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Haoji, Wang, Yiqin, Tang, Yansong, Liu, Yong, Feng, Jiashi, Dai, Jifeng, Jin, Xiaojie |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
par: Zhang, Haoji, et autres
Publié: (2025)
par: Zhang, Haoji, et autres
Publié: (2025)
Hierarchical Memory for Long Video QA
par: Wang, Yiqin, et autres
Publié: (2024)
par: Wang, Yiqin, et autres
Publié: (2024)
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
par: Guo, Zhenghui, et autres
Publié: (2026)
par: Guo, Zhenghui, et autres
Publié: (2026)
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
par: Wang, Yiqin, et autres
Publié: (2024)
par: Wang, Yiqin, et autres
Publié: (2024)
FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
par: Zhuang, Junhao, et autres
Publié: (2025)
par: Zhuang, Junhao, et autres
Publié: (2025)
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
par: Ren, Zhongwei, et autres
Publié: (2026)
par: Ren, Zhongwei, et autres
Publié: (2026)
Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
par: Zhu, Tianrui, et autres
Publié: (2025)
par: Zhu, Tianrui, et autres
Publié: (2025)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
par: Zhang, Haoji, et autres
Publié: (2025)
par: Zhang, Haoji, et autres
Publié: (2025)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
par: Ren, Zhongwei, et autres
Publié: (2025)
par: Ren, Zhongwei, et autres
Publié: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
par: Ma, Fan, et autres
Publié: (2023)
par: Ma, Fan, et autres
Publié: (2023)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
par: Wang, Junxi, et autres
Publié: (2026)
par: Wang, Junxi, et autres
Publié: (2026)
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
par: Lu, Xudong, et autres
Publié: (2026)
par: Lu, Xudong, et autres
Publié: (2026)
VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
par: Wang, Yuji, et autres
Publié: (2025)
par: Wang, Yuji, et autres
Publié: (2025)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
par: Patel, Shrenik, et autres
Publié: (2025)
par: Patel, Shrenik, et autres
Publié: (2025)
FDDet: Achieving Data-Efficient Food Defect Detection Under Real-World Scenarios
par: Xu, Ruihao, et autres
Publié: (2026)
par: Xu, Ruihao, et autres
Publié: (2026)
FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution
par: Chou, Gene, et autres
Publié: (2025)
par: Chou, Gene, et autres
Publié: (2025)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
par: Zuo, Jialong, et autres
Publié: (2025)
par: Zuo, Jialong, et autres
Publié: (2025)
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
par: Bai, Sule, et autres
Publié: (2024)
par: Bai, Sule, et autres
Publié: (2024)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
par: Jin, Hongbo, et autres
Publié: (2025)
par: Jin, Hongbo, et autres
Publié: (2025)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
par: Xu, Ruyi, et autres
Publié: (2025)
par: Xu, Ruyi, et autres
Publié: (2025)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
par: Bai, Sule, et autres
Publié: (2025)
par: Bai, Sule, et autres
Publié: (2025)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
par: Zhang, Gengyuan, et autres
Publié: (2025)
par: Zhang, Gengyuan, et autres
Publié: (2025)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
par: Jin, Xiaojie, et autres
Publié: (2023)
par: Jin, Xiaojie, et autres
Publié: (2023)
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
par: Wang, Zile, et autres
Publié: (2026)
par: Wang, Zile, et autres
Publié: (2026)
Streaming Long Video Understanding with Large Language Models
par: Qian, Rui, et autres
Publié: (2024)
par: Qian, Rui, et autres
Publié: (2024)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
par: Dou, Weijia, et autres
Publié: (2026)
par: Dou, Weijia, et autres
Publié: (2026)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
par: Zhou, Yupeng, et autres
Publié: (2024)
par: Zhou, Yupeng, et autres
Publié: (2024)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
par: Xie, Ming, et autres
Publié: (2026)
par: Xie, Ming, et autres
Publié: (2026)
LOGO: A Long-Form Video Dataset for Group Action Quality Assessment
par: Zhang, Shiyi, et autres
Publié: (2024)
par: Zhang, Shiyi, et autres
Publié: (2024)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
par: Zeng, Xiangyu, et autres
Publié: (2025)
par: Zeng, Xiangyu, et autres
Publié: (2025)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
par: Cheng, Dingxin, et autres
Publié: (2024)
par: Cheng, Dingxin, et autres
Publié: (2024)
StreamSTGS: Streaming Spatial and Temporal Gaussian Grids for Real-Time Free-Viewpoint Video
par: Ke, Zhihui, et autres
Publié: (2025)
par: Ke, Zhihui, et autres
Publié: (2025)
Towards Online Real-Time Memory-based Video Inpainting Transformers
par: Thiry, Guillaume, et autres
Publié: (2024)
par: Thiry, Guillaume, et autres
Publié: (2024)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
par: Song, Yiren, et autres
Publié: (2026)
par: Song, Yiren, et autres
Publié: (2026)
Memory Consolidation Enables Long-Context Video Understanding
par: Balažević, Ivana, et autres
Publié: (2024)
par: Balažević, Ivana, et autres
Publié: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
RIVER: A Real-Time Interaction Benchmark for Video LLMs
par: Shi, Yansong, et autres
Publié: (2026)
par: Shi, Yansong, et autres
Publié: (2026)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
par: Yang, Yanlai, et autres
Publié: (2025)
par: Yang, Yanlai, et autres
Publié: (2025)
Documents similaires
-
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
par: Zhang, Haoji, et autres
Publié: (2025) -
Hierarchical Memory for Long Video QA
par: Wang, Yiqin, et autres
Publié: (2024) -
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
par: Guo, Zhenghui, et autres
Publié: (2026) -
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
par: Wang, Yiqin, et autres
Publié: (2024) -
FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
par: Zhuang, Junhao, et autres
Publié: (2025)