Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Junxiao, Wang, Jingjing, Luo, Jiamin, Yu, Peiying, Zhou, Guodong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
by: Luo, Jiamin, et al.
Published: (2025)
by: Luo, Jiamin, et al.
Published: (2025)
Towards LLM-centric Affective Visual Customization via Efficient and Precise Emotion Manipulating
by: Luo, Jiamin, et al.
Published: (2026)
by: Luo, Jiamin, et al.
Published: (2026)
DeepSVU: Towards In-depth Security-oriented Video Understanding via Unified Physical-world Regularized MoE
by: Jin, Yujie, et al.
Published: (2026)
by: Jin, Yujie, et al.
Published: (2026)
RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction
by: Jin, Yongkang, et al.
Published: (2026)
by: Jin, Yongkang, et al.
Published: (2026)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)
by: Lin, Han, et al.
Published: (2023)
DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video
by: Wen, Hao, et al.
Published: (2025)
by: Wen, Hao, et al.
Published: (2025)
TopicDiff: A Topic-enriched Diffusion Approach for Multimodal Conversational Emotion Detection
by: Luo, Jiamin, et al.
Published: (2024)
by: Luo, Jiamin, et al.
Published: (2024)
Towards Open-Vocabulary Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
Abnormal Event Detection In Videos Using Deep Embedding
by: Venkatrayappa, Darshan
Published: (2024)
by: Venkatrayappa, Darshan
Published: (2024)
From Sim-to-Real: Toward General Event-based Low-light Frame Interpolation with Per-scene Optimization
by: Zhang, Ziran, et al.
Published: (2024)
by: Zhang, Ziran, et al.
Published: (2024)
ChatASU: Evoking LLM's Reflexion to Truly Understand Aspect Sentiment in Dialogues
by: Liu, Yiding, et al.
Published: (2024)
by: Liu, Yiding, et al.
Published: (2024)
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
by: Liu, Yunze, et al.
Published: (2026)
by: Liu, Yunze, et al.
Published: (2026)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution
by: Li, Wenxue, et al.
Published: (2026)
by: Li, Wenxue, et al.
Published: (2026)
Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
by: Yu, Fei, et al.
Published: (2024)
by: Yu, Fei, et al.
Published: (2024)
Recognition of Abnormal Events in Surveillance Videos using Weakly Supervised Dual-Encoder Models
by: Tsfaty, Noam, et al.
Published: (2025)
by: Tsfaty, Noam, et al.
Published: (2025)
Localizing Events in Videos with Multimodal Queries
by: Zhang, Gengyuan, et al.
Published: (2024)
by: Zhang, Gengyuan, et al.
Published: (2024)
How to Understand "Support"? An Implicit-enhanced Causal Inference Approach for Weakly-supervised Phrase Grounding
by: Luo, Jiamin, et al.
Published: (2024)
by: Luo, Jiamin, et al.
Published: (2024)
Local-Global Context Aware Transformer for Language-Guided Video Segmentation
by: Liang, Chen, et al.
Published: (2022)
by: Liang, Chen, et al.
Published: (2022)
ECHO: Event-Centric Hypergraph Operations via Multi-Agent Collaboration for Multimedia Event Extraction
by: Chu, Hailong, et al.
Published: (2026)
by: Chu, Hailong, et al.
Published: (2026)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
by: Li, Yian, et al.
Published: (2026)
by: Li, Yian, et al.
Published: (2026)
Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
by: Chen, Yuxiao, et al.
Published: (2024)
by: Chen, Yuxiao, et al.
Published: (2024)
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
by: Ma, Yongjia, et al.
Published: (2025)
by: Ma, Yongjia, et al.
Published: (2025)
Cambrian-S: Towards Spatial Supersensing in Video
by: Yang, Shusheng, et al.
Published: (2025)
by: Yang, Shusheng, et al.
Published: (2025)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
Multi-Class Abnormality Classification Task in Video Capsule Endoscopy
by: Verma, Dev Rishi, et al.
Published: (2024)
by: Verma, Dev Rishi, et al.
Published: (2024)
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
by: Yu, Fangxu, et al.
Published: (2026)
by: Yu, Fangxu, et al.
Published: (2026)
Learning Global and Local Features of Normal Brain Anatomy for Unsupervised Abnormality Detection
by: Kobayashi, Kazuma, et al.
Published: (2020)
by: Kobayashi, Kazuma, et al.
Published: (2020)
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
by: Mou, Tingshu, et al.
Published: (2026)
by: Mou, Tingshu, et al.
Published: (2026)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
Towards Event-oriented Long Video Understanding
by: Du, Yifan, et al.
Published: (2024)
by: Du, Yifan, et al.
Published: (2024)
HumanMM: Global Human Motion Recovery from Multi-shot Videos
by: Zhang, Yuhong, et al.
Published: (2025)
by: Zhang, Yuhong, et al.
Published: (2025)
Sherlock: Self-Correcting Reasoning in Vision-Language Models
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
Table-Critic: A Multi-Agent Framework for Collaborative Criticism and Refinement in Table Reasoning
by: Yu, Peiying, et al.
Published: (2025)
by: Yu, Peiying, et al.
Published: (2025)
3D scene generation from scene graphs and self-attention
by: Bonazzi, Pietro, et al.
Published: (2024)
by: Bonazzi, Pietro, et al.
Published: (2024)
Similar Items
-
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
by: Luo, Jiamin, et al.
Published: (2025) -
Towards LLM-centric Affective Visual Customization via Efficient and Precise Emotion Manipulating
by: Luo, Jiamin, et al.
Published: (2026) -
DeepSVU: Towards In-depth Security-oriented Video Understanding via Unified Physical-world Regularized MoE
by: Jin, Yujie, et al.
Published: (2026) -
RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction
by: Jin, Yongkang, et al.
Published: (2026) -
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)