OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Zhijia, Li, Jiaming, Chen, Weikai, Zhang, Yanhao, Lu, Haonan, Li, Guanbin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2026)
by: Li, Jiaming, et al.
Published: (2026)
Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
by: Wang, Kuo, et al.
Published: (2025)
by: Wang, Kuo, et al.
Published: (2025)
OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd Representation
by: Zhao, Ganlong, et al.
Published: (2024)
by: Zhao, Ganlong, et al.
Published: (2024)
Thinking in Streaming Video
by: Liu, Zikang, et al.
Published: (2026)
by: Liu, Zikang, et al.
Published: (2026)
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
by: Song, Xinshuai, et al.
Published: (2024)
by: Song, Xinshuai, et al.
Published: (2024)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
by: Zeng, Xiangyu, et al.
Published: (2025)
by: Zeng, Xiangyu, et al.
Published: (2025)
DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery Analysis
by: Cai, Yinqi, et al.
Published: (2025)
by: Cai, Yinqi, et al.
Published: (2025)
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
Click-to-Ask: An AI Live Streaming Assistant with Offline Copywriting and Online Interactive QA
by: Yu, Ruizhi, et al.
Published: (2026)
by: Yu, Ruizhi, et al.
Published: (2026)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
by: Wang, Kuo, et al.
Published: (2024)
by: Wang, Kuo, et al.
Published: (2024)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
by: Zhang, Gengyuan, et al.
Published: (2025)
by: Zhang, Gengyuan, et al.
Published: (2025)
AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
by: Zhang, Ruifei, et al.
Published: (2025)
by: Zhang, Ruifei, et al.
Published: (2025)
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2024)
by: Li, Jiaming, et al.
Published: (2024)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
by: Cheng, Dingxin, et al.
Published: (2024)
by: Cheng, Dingxin, et al.
Published: (2024)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
by: Wen, Siwei, et al.
Published: (2026)
by: Wen, Siwei, et al.
Published: (2026)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
by: Wang, Junxi, et al.
Published: (2026)
by: Wang, Junxi, et al.
Published: (2026)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
by: Xie, Yiweng, et al.
Published: (2026)
by: Xie, Yiweng, et al.
Published: (2026)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
by: Zheng, Minghang, et al.
Published: (2025)
by: Zheng, Minghang, et al.
Published: (2025)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
by: Zhang, Haowei, et al.
Published: (2026)
by: Zhang, Haowei, et al.
Published: (2026)
HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
DeformStream: Deformation-based Adaptive Volumetric Video Streaming
by: Li, Boyan, et al.
Published: (2024)
by: Li, Boyan, et al.
Published: (2024)
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
by: Wu, Qi, et al.
Published: (2025)
by: Wu, Qi, et al.
Published: (2025)
Hierarchical Memory for Long Video QA
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
by: Zhong, Chunlin, et al.
Published: (2025)
by: Zhong, Chunlin, et al.
Published: (2025)
3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
by: Wei, Zeming, et al.
Published: (2025)
by: Wei, Zeming, et al.
Published: (2025)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
by: He, Zijian, et al.
Published: (2024)
by: He, Zijian, et al.
Published: (2024)
Adaptive Event Stream Slicing for Open-Vocabulary Event-Based Object Detection via Vision-Language Knowledge Distillation
by: Zhang, Jinchang, et al.
Published: (2025)
by: Zhang, Jinchang, et al.
Published: (2025)
EvSign: Sign Language Recognition and Translation with Streaming Events
by: Zhang, Pengyu, et al.
Published: (2024)
by: Zhang, Pengyu, et al.
Published: (2024)
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
by: Xie, Qingsong, et al.
Published: (2025)
by: Xie, Qingsong, et al.
Published: (2025)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
by: Liu, Ruiping, et al.
Published: (2026)
by: Liu, Ruiping, et al.
Published: (2026)
Compositional Physical Reasoning of Objects and Events from Videos
by: Chen, Zhenfang, et al.
Published: (2024)
by: Chen, Zhenfang, et al.
Published: (2024)
EventTracer: Fast Path Tracing-based Event Stream Rendering
by: Li, Zhenyang, et al.
Published: (2025)
by: Li, Zhenyang, et al.
Published: (2025)
DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
by: Li, Weihong, et al.
Published: (2025)
by: Li, Weihong, et al.
Published: (2025)
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
by: Liao, Zhenyi, et al.
Published: (2025)
by: Liao, Zhenyi, et al.
Published: (2025)
Zebrafish Counting Using Event Stream Data
by: Chen, Qianghua, et al.
Published: (2025)
by: Chen, Qianghua, et al.
Published: (2025)
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
by: Wang, Yiyu, et al.
Published: (2025)
by: Wang, Yiyu, et al.
Published: (2025)
Similar Items
-
GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2026) -
Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
by: Wang, Kuo, et al.
Published: (2025) -
OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd Representation
by: Zhao, Ganlong, et al.
Published: (2024) -
Thinking in Streaming Video
by: Liu, Zikang, et al.
Published: (2026) -
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
by: Song, Xinshuai, et al.
Published: (2024)