PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Erregue, Iñaki, Nasrollahi, Kamal, Escalera, Sergio |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
YOLO11-JDE: Fast and Accurate Multi-Object Tracking with Self-Supervised Re-ID
by: Erregue, Iñaki, et al.
Published: (2025)
by: Erregue, Iñaki, et al.
Published: (2025)
Machine Unlearning in Hyperbolic vs. Euclidean Multimodal Contrastive Learning: Adapting Alignment Calibration to MERU
by: Vidal, Àlex Pujol, et al.
Published: (2025)
by: Vidal, Àlex Pujol, et al.
Published: (2025)
SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models
by: Rabasseda, Oriol, et al.
Published: (2026)
by: Rabasseda, Oriol, et al.
Published: (2026)
Balancing Privacy and Action Performance: A Penalty-Driven Approach to Image Anonymization
by: Aslam, Nazia, et al.
Published: (2025)
by: Aslam, Nazia, et al.
Published: (2025)
VideoPrism: A Foundational Visual Encoder for Video Understanding
by: Zhao, Long, et al.
Published: (2024)
by: Zhao, Long, et al.
Published: (2024)
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning
by: Zhu, Liyun, et al.
Published: (2025)
by: Zhu, Liyun, et al.
Published: (2025)
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
by: Pereira, João, et al.
Published: (2026)
by: Pereira, João, et al.
Published: (2026)
Video Anomaly Detection with Contours -- A Study
by: Siemon, Mia, et al.
Published: (2025)
by: Siemon, Mia, et al.
Published: (2025)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
by: Lin, Rui, et al.
Published: (2026)
by: Lin, Rui, et al.
Published: (2026)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
Bounding Boxes and Probabilistic Graphical Models: Video Anomaly Detection Simplified
by: Siemon, Mia, et al.
Published: (2024)
by: Siemon, Mia, et al.
Published: (2024)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
by: Guo, Bangwei, et al.
Published: (2025)
by: Guo, Bangwei, et al.
Published: (2025)
Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
by: Wu, Peng, et al.
Published: (2024)
by: Wu, Peng, et al.
Published: (2024)
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
by: Wang, Yufei, et al.
Published: (2025)
by: Wang, Yufei, et al.
Published: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
by: Zou, Shu, et al.
Published: (2025)
by: Zou, Shu, et al.
Published: (2025)
Retrieval, Refinement, and Ranking for Text-to-Video Generation via Prompt Optimization and Test-Time Scaling
by: Rahman, Zillur, et al.
Published: (2026)
by: Rahman, Zillur, et al.
Published: (2026)
Apollo: An Exploration of Video Understanding in Large Multimodal Models
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
by: Zhao, Xiran, et al.
Published: (2026)
by: Zhao, Xiran, et al.
Published: (2026)
Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance
by: Sun, Shengyang, et al.
Published: (2026)
by: Sun, Shengyang, et al.
Published: (2026)
CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
by: Bhosale, Mahesh, et al.
Published: (2026)
by: Bhosale, Mahesh, et al.
Published: (2026)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
by: Zhou, Qiji, et al.
Published: (2024)
by: Zhou, Qiji, et al.
Published: (2024)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
by: Zhang, Zhihong, et al.
Published: (2025)
by: Zhang, Zhihong, et al.
Published: (2025)
CL-MAE: Curriculum-Learned Masked Autoencoders
by: Madan, Neelu, et al.
Published: (2023)
by: Madan, Neelu, et al.
Published: (2023)
Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
by: Lee, Yujin, et al.
Published: (2024)
by: Lee, Yujin, et al.
Published: (2024)
Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly
by: Du, Hang, et al.
Published: (2024)
by: Du, Hang, et al.
Published: (2024)
Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly
by: Du, Hang, et al.
Published: (2024)
by: Du, Hang, et al.
Published: (2024)
Evaluating the Effectiveness of Video Anomaly Detection in the Wild: Online Learning and Inference for Real-world Deployment
by: Yao, Shanle, et al.
Published: (2024)
by: Yao, Shanle, et al.
Published: (2024)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
by: Jang, Lawrence, et al.
Published: (2024)
by: Jang, Lawrence, et al.
Published: (2024)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
by: Peng, Tianhao, et al.
Published: (2025)
by: Peng, Tianhao, et al.
Published: (2025)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
by: Wu, Peng, et al.
Published: (2023)
by: Wu, Peng, et al.
Published: (2023)
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
by: Yakun, Cui, et al.
Published: (2025)
by: Yakun, Cui, et al.
Published: (2025)
Prism: Spectral-Aware Block-Sparse Attention
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
Injecting Explainability and Lightweight Design into Weakly Supervised Video Anomaly Detection Systems
by: Jiang, Wen-Dong, et al.
Published: (2024)
by: Jiang, Wen-Dong, et al.
Published: (2024)
Similar Items
-
YOLO11-JDE: Fast and Accurate Multi-Object Tracking with Self-Supervised Re-ID
by: Erregue, Iñaki, et al.
Published: (2025) -
Machine Unlearning in Hyperbolic vs. Euclidean Multimodal Contrastive Learning: Adapting Alignment Calibration to MERU
by: Vidal, Àlex Pujol, et al.
Published: (2025) -
SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models
by: Rabasseda, Oriol, et al.
Published: (2026) -
Balancing Privacy and Action Performance: A Penalty-Driven Approach to Image Anonymization
by: Aslam, Nazia, et al.
Published: (2025) -
VideoPrism: A Foundational Visual Encoder for Video Understanding
by: Zhao, Long, et al.
Published: (2024)