Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Gengyuan, Ding, Mingcong, Liu, Tong, Zhang, Yao, Tresp, Volker |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
von: Liao, Ruotong, et al.
Veröffentlicht: (2024)
von: Liao, Ruotong, et al.
Veröffentlicht: (2024)
Multi-event Video-Text Retrieval
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
von: Amoroso, Roberto, et al.
Veröffentlicht: (2024)
von: Amoroso, Roberto, et al.
Veröffentlicht: (2024)
Localizing Events in Videos with Multimodal Queries
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2024)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
von: Lin, Junming, et al.
Veröffentlicht: (2024)
von: Lin, Junming, et al.
Veröffentlicht: (2024)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
von: Sun, Peiwen, et al.
Veröffentlicht: (2026)
von: Sun, Peiwen, et al.
Veröffentlicht: (2026)
When and Where do Events Switch in Multi-Event Video Generation?
von: Liao, Ruotong, et al.
Veröffentlicht: (2025)
von: Liao, Ruotong, et al.
Veröffentlicht: (2025)
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
von: Liang, Zhijia, et al.
Veröffentlicht: (2026)
von: Liang, Zhijia, et al.
Veröffentlicht: (2026)
Mixup Helps Understanding Multimodal Video Better
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2025)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2024)
von: Zhang, Haoji, et al.
Veröffentlicht: (2024)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2025)
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)
EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
von: Liu, Shaoyu, et al.
Veröffentlicht: (2025)
von: Liu, Shaoyu, et al.
Veröffentlicht: (2025)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
EventGPT: Event Stream Understanding with Multimodal Large Language Models
von: Liu, Shaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Shaoyu, et al.
Veröffentlicht: (2024)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
von: Wen, Siwei, et al.
Veröffentlicht: (2026)
von: Wen, Siwei, et al.
Veröffentlicht: (2026)
EventFlash: Towards Efficient MLLMs for Event-Based Vision
von: Liu, Shaoyu, et al.
Veröffentlicht: (2026)
von: Liu, Shaoyu, et al.
Veröffentlicht: (2026)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
von: Xie, Ming, et al.
Veröffentlicht: (2026)
von: Xie, Ming, et al.
Veröffentlicht: (2026)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
Multimodal Pragmatic Jailbreak on Text-to-image Models
von: Liu, Tong, et al.
Veröffentlicht: (2024)
von: Liu, Tong, et al.
Veröffentlicht: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
von: He, Yuping, et al.
Veröffentlicht: (2025)
von: He, Yuping, et al.
Veröffentlicht: (2025)
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
von: Liu, Xianjie, et al.
Veröffentlicht: (2026)
von: Liu, Xianjie, et al.
Veröffentlicht: (2026)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
von: Chen, Jian, et al.
Veröffentlicht: (2024)
von: Chen, Jian, et al.
Veröffentlicht: (2024)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026)
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
von: Pang, Zhanzhong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025) -
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
von: Liao, Ruotong, et al.
Veröffentlicht: (2024) -
Multi-event Video-Text Retrieval
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023) -
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023) -
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
von: Amoroso, Roberto, et al.
Veröffentlicht: (2024)