ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Zihao, Wu, Xiaoyu, Li, Wenna, Wu, Jianqin, Yang, Linlin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Language-guided Open-world Video Anomaly Detection under Weak Supervision
par: Liu, Zihao, et autres
Publié: (2025)
par: Liu, Zihao, et autres
Publié: (2025)
Rethinking Metrics and Benchmarks of Video Anomaly Detection
par: Liu, Zihao, et autres
Publié: (2025)
par: Liu, Zihao, et autres
Publié: (2025)
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
par: Zhu, Zhiyi, et autres
Publié: (2025)
par: Zhu, Zhiyi, et autres
Publié: (2025)
Hawk: Learning to Understand Open-World Video Anomalies
par: Tang, Jiaqi, et autres
Publié: (2024)
par: Tang, Jiaqi, et autres
Publié: (2024)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
par: Lin, Junming, et autres
Publié: (2024)
par: Lin, Junming, et autres
Publié: (2024)
An Efficient Streaming Video Understanding Framework with Agentic Control
par: Liu, Jinming, et autres
Publié: (2026)
par: Liu, Jinming, et autres
Publié: (2026)
Open-Vocabulary Video Anomaly Detection
par: Wu, Peng, et autres
Publié: (2023)
par: Wu, Peng, et autres
Publié: (2023)
Learning Prompt-Enhanced Context Features for Weakly-Supervised Video Anomaly Detection
par: Pu, Yujiang, et autres
Publié: (2023)
par: Pu, Yujiang, et autres
Publié: (2023)
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
par: Jiang, Yuanyuan, et autres
Publié: (2024)
par: Jiang, Yuanyuan, et autres
Publié: (2024)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
par: Gao, Shibo, et autres
Publié: (2025)
par: Gao, Shibo, et autres
Publié: (2025)
SVL: Spike-based Vision-language Pretraining for Efficient 3D Open-world Understanding
par: Qiu, Xuerui, et autres
Publié: (2025)
par: Qiu, Xuerui, et autres
Publié: (2025)
SRVAU-R1: Enhancing Video Anomaly Understanding via Reflection-Aware Learning
par: Zhao, Zihao, et autres
Publié: (2026)
par: Zhao, Zihao, et autres
Publié: (2026)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
par: Wang, Junxi, et autres
Publié: (2026)
par: Wang, Junxi, et autres
Publié: (2026)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
par: Kang, Hyolim, et autres
Publié: (2025)
par: Kang, Hyolim, et autres
Publié: (2025)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
par: Zhang, Haoji, et autres
Publié: (2025)
par: Zhang, Haoji, et autres
Publié: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
par: Zeng, Xiangyu, et autres
Publié: (2025)
par: Zeng, Xiangyu, et autres
Publié: (2025)
VideoScan: Enabling Efficient Streaming Video Understanding via Frame-level Semantic Carriers
par: Li, Ruanjun, et autres
Publié: (2025)
par: Li, Ruanjun, et autres
Publié: (2025)
Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
par: Zhang, Yulin, et autres
Publié: (2025)
par: Zhang, Yulin, et autres
Publié: (2025)
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
par: Yang, Zhenyu, et autres
Publié: (2025)
par: Yang, Zhenyu, et autres
Publié: (2025)
CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
par: Yu, Yating, et autres
Publié: (2025)
par: Yu, Yating, et autres
Publié: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
par: Chen, Xueyi, et autres
Publié: (2025)
par: Chen, Xueyi, et autres
Publié: (2025)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
par: Wang, Yifei, et autres
Publié: (2025)
par: Wang, Yifei, et autres
Publié: (2025)
EA3D: Online Open-World 3D Object Extraction from Streaming Videos
par: Zhou, Xiaoyu, et autres
Publié: (2025)
par: Zhou, Xiaoyu, et autres
Publié: (2025)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
par: Sun, Peiwen, et autres
Publié: (2026)
par: Sun, Peiwen, et autres
Publié: (2026)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
par: Yang, Haolin, et autres
Publié: (2025)
par: Yang, Haolin, et autres
Publié: (2025)
E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
par: Liu, Ye, et autres
Publié: (2024)
par: Liu, Ye, et autres
Publié: (2024)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
par: Ning, Zhenyu, et autres
Publié: (2025)
par: Ning, Zhenyu, et autres
Publié: (2025)
Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets
par: Wu, Peng, et autres
Publié: (2026)
par: Wu, Peng, et autres
Publié: (2026)
Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly Detection
par: Yang, Zhiwei, et autres
Publié: (2024)
par: Yang, Zhiwei, et autres
Publié: (2024)
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
par: Lu, Xudong, et autres
Publié: (2026)
par: Lu, Xudong, et autres
Publié: (2026)
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
par: Xie, Ming, et autres
Publié: (2026)
par: Xie, Ming, et autres
Publié: (2026)
Anomize: Better Open Vocabulary Video Anomaly Detection
par: Li, Fei, et autres
Publié: (2025)
par: Li, Fei, et autres
Publié: (2025)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
par: Chen, Feng, et autres
Publié: (2024)
par: Chen, Feng, et autres
Publié: (2024)
The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
par: Wu, Zhuoyuan, et autres
Publié: (2025)
par: Wu, Zhuoyuan, et autres
Publié: (2025)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
par: Wu, Hang, et autres
Publié: (2026)
par: Wu, Hang, et autres
Publié: (2026)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
par: Wu, Peiran, et autres
Publié: (2025)
par: Wu, Peiran, et autres
Publié: (2025)
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
par: Wu, Peiran, et autres
Publié: (2026)
par: Wu, Peiran, et autres
Publié: (2026)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
par: Patel, Shrenik, et autres
Publié: (2025)
par: Patel, Shrenik, et autres
Publié: (2025)
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
par: Luo, Yawen, et autres
Publié: (2026)
par: Luo, Yawen, et autres
Publié: (2026)
Documents similaires
-
Language-guided Open-world Video Anomaly Detection under Weak Supervision
par: Liu, Zihao, et autres
Publié: (2025) -
Rethinking Metrics and Benchmarks of Video Anomaly Detection
par: Liu, Zihao, et autres
Publié: (2025) -
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
par: Zhu, Zhiyi, et autres
Publié: (2025) -
Hawk: Learning to Understand Open-World Video Anomalies
par: Tang, Jiaqi, et autres
Publié: (2024) -
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
par: Lin, Junming, et autres
Publié: (2024)