A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lin, Dongheng, Qu, Mengxue, Han, Kunyang, Jiao, Jianbo, Jin, Xiaojie, Wei, Yunchao |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images
par: Lin, Dongheng, et autres
Publié: (2025)
par: Lin, Dongheng, et autres
Publié: (2025)
ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model
par: Han, Kunyang, et autres
Publié: (2024)
par: Han, Kunyang, et autres
Publié: (2024)
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
par: Kang, Weitai, et autres
Publié: (2024)
par: Kang, Weitai, et autres
Publié: (2024)
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
par: Zhao, Shifang, et autres
Publié: (2025)
par: Zhao, Shifang, et autres
Publié: (2025)
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
par: Ahn, Sunghyun, et autres
Publié: (2025)
par: Ahn, Sunghyun, et autres
Publié: (2025)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
par: Ren, Zhongwei, et autres
Publié: (2025)
par: Ren, Zhongwei, et autres
Publié: (2025)
Lane2Seq: Towards Unified Lane Detection via Sequence Generation
par: Zhou, Kunyang
Publié: (2024)
par: Zhou, Kunyang
Publié: (2024)
Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection
par: Qu, Zhen, et autres
Publié: (2025)
par: Qu, Zhen, et autres
Publié: (2025)
Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
par: Kang, Weitai, et autres
Publié: (2024)
par: Kang, Weitai, et autres
Publié: (2024)
3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation
par: Chen, Wenxin, et autres
Publié: (2025)
par: Chen, Wenxin, et autres
Publié: (2025)
PixelLM: Pixel Reasoning with Large Multimodal Model
par: Ren, Zhongwei, et autres
Publié: (2023)
par: Ren, Zhongwei, et autres
Publié: (2023)
AnomalyPainter: Vision-Language-Diffusion Synergy for Zero-Shot Realistic and Diverse Industrial Anomaly Synthesis
par: Lai, Zhangyu, et autres
Publié: (2025)
par: Lai, Zhangyu, et autres
Publié: (2025)
Bridge the Points: Graph-based Few-shot Segment Anything Semantically
par: Zhang, Anqi, et autres
Publié: (2024)
par: Zhang, Anqi, et autres
Publié: (2024)
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
par: Ren, Zhongwei, et autres
Publié: (2026)
par: Ren, Zhongwei, et autres
Publié: (2026)
CoPS: Conditional Prompt Synthesis for Zero-Shot Anomaly Detection
par: Chen, Qiyu, et autres
Publié: (2025)
par: Chen, Qiyu, et autres
Publié: (2025)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
par: Li, Kunyang, et autres
Publié: (2026)
par: Li, Kunyang, et autres
Publié: (2026)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
par: Yu, Xiao, et autres
Publié: (2025)
par: Yu, Xiao, et autres
Publié: (2025)
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
par: Meng, Yihao, et autres
Publié: (2025)
par: Meng, Yihao, et autres
Publié: (2025)
Bootstrap Fine-Grained Vision-Language Alignment for Unified Zero-Shot Anomaly Localization
par: Deng, Hanqiu, et autres
Publié: (2023)
par: Deng, Hanqiu, et autres
Publié: (2023)
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
par: Liu, Man, et autres
Publié: (2024)
par: Liu, Man, et autres
Publié: (2024)
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
par: Dai, Zunkai, et autres
Publié: (2026)
par: Dai, Zunkai, et autres
Publié: (2026)
Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
par: Li, Kaiqiang, et autres
Publié: (2026)
par: Li, Kaiqiang, et autres
Publié: (2026)
CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental Segmentation
par: Fang, Kai, et autres
Publié: (2025)
par: Fang, Kai, et autres
Publié: (2025)
Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token
par: Zhang, Anqi, et autres
Publié: (2026)
par: Zhang, Anqi, et autres
Publié: (2026)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
par: Jin, Woojeong, et autres
Publié: (2026)
par: Jin, Woojeong, et autres
Publié: (2026)
Exploring Image Representation with Decoupled Classical Visual Descriptors
par: Qu, Chenyuan, et autres
Publié: (2025)
par: Qu, Chenyuan, et autres
Publié: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
par: Fang, Xinyu, et autres
Publié: (2024)
par: Fang, Xinyu, et autres
Publié: (2024)
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models
par: Xu, Jiacong, et autres
Publié: (2025)
par: Xu, Jiacong, et autres
Publié: (2025)
Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
par: Yin, Xiaojie, et autres
Publié: (2025)
par: Yin, Xiaojie, et autres
Publié: (2025)
Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
par: Lai, Yuxiang, et autres
Publié: (2025)
par: Lai, Yuxiang, et autres
Publié: (2025)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
par: Kao, Shiu-hong, et autres
Publié: (2025)
par: Kao, Shiu-hong, et autres
Publié: (2025)
The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
par: Gao, Shibo, et autres
Publié: (2025)
par: Gao, Shibo, et autres
Publié: (2025)
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
par: Tan, Chuangchuang, et autres
Publié: (2025)
par: Tan, Chuangchuang, et autres
Publié: (2025)
DiffIR2VR-Zero: Zero-Shot Video Restoration with Diffusion-based Image Restoration Models
par: Yeh, Chang-Han, et autres
Publié: (2024)
par: Yeh, Chang-Han, et autres
Publié: (2024)
MAGIC: Map-Guided Few-Shot Audio-Visual Acoustics Modeling
par: Huang, Diwei, et autres
Publié: (2024)
par: Huang, Diwei, et autres
Publié: (2024)
ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection
par: Chen, Qiuhui, et autres
Publié: (2026)
par: Chen, Qiuhui, et autres
Publié: (2026)
Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval
par: Alavi, Ali
Publié: (2026)
par: Alavi, Ali
Publié: (2026)
On the Problem of Consistent Anomalies in Zero-Shot Industrial Anomaly Detection
par: Le-Gia, Tai, et autres
Publié: (2025)
par: Le-Gia, Tai, et autres
Publié: (2025)
Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generation
par: Jia, Jun, et autres
Publié: (2025)
par: Jia, Jun, et autres
Publié: (2025)
PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
par: Huang, Po-Han, et autres
Publié: (2025)
par: Huang, Po-Han, et autres
Publié: (2025)
Documents similaires
-
What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images
par: Lin, Dongheng, et autres
Publié: (2025) -
ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model
par: Han, Kunyang, et autres
Publié: (2024) -
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
par: Kang, Weitai, et autres
Publié: (2024) -
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
par: Zhao, Shifang, et autres
Publié: (2025) -
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
par: Ahn, Sunghyun, et autres
Publié: (2025)