Multi-turn Physics-informed Vision-language Model for Physics-grounded Anomaly Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Yao, Xu, Xiaohao, Wu, Yingna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
by: Li, Wenqiao, et al.
Published: (2025)
by: Li, Wenqiao, et al.
Published: (2025)
Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided Alignment
by: Chen, Xintao, et al.
Published: (2025)
by: Chen, Xintao, et al.
Published: (2025)
Breaking the Rigid Prior: Towards Articulated 3D Anomaly Detection
by: Gan, Jinye, et al.
Published: (2026)
by: Gan, Jinye, et al.
Published: (2026)
Multi-Sensor Object Anomaly Detection: Unifying Appearance, Geometry, and Internal Properties
by: Li, Wenqiao, et al.
Published: (2024)
by: Li, Wenqiao, et al.
Published: (2024)
Bridging 3D Anomaly Localization and Repair via High-Quality Continuous Geometric Representation
by: Zheng, Bozhong, et al.
Published: (2025)
by: Zheng, Bozhong, et al.
Published: (2025)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Complementary Pseudo Multimodal Feature for Point Cloud Anomaly Detection
by: Cao, Yunkang, et al.
Published: (2023)
by: Cao, Yunkang, et al.
Published: (2023)
Learn Suspected Anomalies from Event Prompts for Video Anomaly Detection
by: Tao, Chenchen, et al.
Published: (2024)
by: Tao, Chenchen, et al.
Published: (2024)
MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
by: Zhao, Xiran, et al.
Published: (2026)
by: Zhao, Xiran, et al.
Published: (2026)
LogiCode: an LLM-Driven Framework for Logical Anomaly Detection
by: Zhang, Yiheng, et al.
Published: (2024)
by: Zhang, Yiheng, et al.
Published: (2024)
Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
Photorealistic Phantom Roads in Real Scenes: Disentangling 3D Hallucinations from Physical Geometry
by: Nguyen, Hoang, et al.
Published: (2025)
by: Nguyen, Hoang, et al.
Published: (2025)
VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
by: Hicsonmez, Samet, et al.
Published: (2025)
by: Hicsonmez, Samet, et al.
Published: (2025)
Towards Physics-informed Diffusion for Anomaly Detection in Trajectories
by: Sharma, Arun, et al.
Published: (2025)
by: Sharma, Arun, et al.
Published: (2025)
GlanceVAD: Exploring Glance Supervision for Label-efficient Video Anomaly Detection
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
A Survey on Visual Anomaly Detection: Challenge, Approach, and Prospect
by: Cao, Yunkang, et al.
Published: (2024)
by: Cao, Yunkang, et al.
Published: (2024)
Visual Anomaly Detection under Complex View-Illumination Interplay: A Large-Scale Benchmark
by: Cao, Yunkang, et al.
Published: (2025)
by: Cao, Yunkang, et al.
Published: (2025)
Center-aware Residual Anomaly Synthesis for Multi-class Industrial Anomaly Detection
by: Chen, Qiyu, et al.
Published: (2025)
by: Chen, Qiyu, et al.
Published: (2025)
Topo-R1: Detecting Topological Anomalies via Vision-Language Models
by: Xu, Meilong, et al.
Published: (2026)
by: Xu, Meilong, et al.
Published: (2026)
Myriad: Large Multimodal Model by Applying Vision Experts for Industrial Anomaly Detection
by: Li, Yuanze, et al.
Published: (2023)
by: Li, Yuanze, et al.
Published: (2023)
Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
by: Zhou, Lihua, et al.
Published: (2025)
by: Zhou, Lihua, et al.
Published: (2025)
Vision-Language Models Assisted Unsupervised Video Anomaly Detection
by: Jiang, Yalong, et al.
Published: (2024)
by: Jiang, Yalong, et al.
Published: (2024)
Detect Closer Surfaces that can be Seen: New Modeling and Evaluation in Cross-domain 3D Object Detection
by: Zhang, Ruixiao, et al.
Published: (2024)
by: Zhang, Ruixiao, et al.
Published: (2024)
GV-VAD : Exploring Video Generation for Weakly-Supervised Video Anomaly Detection
by: Cai, Suhang, et al.
Published: (2025)
by: Cai, Suhang, et al.
Published: (2025)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
Road Rage Reasoning with Vision-language Models (VLMs): Task Definition and Evaluation Dataset
by: Weng, Yibing, et al.
Published: (2025)
by: Weng, Yibing, et al.
Published: (2025)
Knowledge-grounded Adaptation Strategy for Vision-language Models: Building Unique Case-set for Screening Mammograms for Residents Training
by: Khan, Aisha Urooj, et al.
Published: (2024)
by: Khan, Aisha Urooj, et al.
Published: (2024)
Towards Training-free Anomaly Detection with Vision and Language Foundation Models
by: Zhang, Jinjin, et al.
Published: (2025)
by: Zhang, Jinjin, et al.
Published: (2025)
Anomaly Detection by Adapting a pre-trained Vision Language Model
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
CLIP3D-AD: Extending CLIP for 3D Few-Shot Anomaly Detection with Multi-View Images Generation
by: Zuo, Zuo, et al.
Published: (2024)
by: Zuo, Zuo, et al.
Published: (2024)
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
by: Zhang, Xing, et al.
Published: (2024)
by: Zhang, Xing, et al.
Published: (2024)
Frequency-Guided Multi-Level Human Action Anomaly Detection with Normalizing Flows
by: Maeda, Shun, et al.
Published: (2024)
by: Maeda, Shun, et al.
Published: (2024)
In-context Prompt Learning for Test-time Vision Recognition with Frozen Vision-language Model
by: Yin, Junhui, et al.
Published: (2024)
by: Yin, Junhui, et al.
Published: (2024)
One Language-Free Foundation Model Is Enough for Universal Vision Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2026)
by: Gao, Bin-Bin, et al.
Published: (2026)
Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection
by: Qian, Kun, et al.
Published: (2024)
by: Qian, Kun, et al.
Published: (2024)
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
by: Dong, Jihao, et al.
Published: (2024)
by: Dong, Jihao, et al.
Published: (2024)
Multi-turn Consistent Image Editing
by: Zhou, Zijun, et al.
Published: (2025)
by: Zhou, Zijun, et al.
Published: (2025)
AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
by: Gu, Zhaopeng, et al.
Published: (2025)
by: Gu, Zhaopeng, et al.
Published: (2025)
ADPretrain: Advancing Industrial Anomaly Detection via Anomaly Representation Pretraining
by: Yao, Xincheng, et al.
Published: (2025)
by: Yao, Xincheng, et al.
Published: (2025)
VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction
by: Nakata, Hiroto, et al.
Published: (2026)
by: Nakata, Hiroto, et al.
Published: (2026)
Similar Items
-
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
by: Li, Wenqiao, et al.
Published: (2025) -
Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided Alignment
by: Chen, Xintao, et al.
Published: (2025) -
Breaking the Rigid Prior: Towards Articulated 3D Anomaly Detection
by: Gan, Jinye, et al.
Published: (2026) -
Multi-Sensor Object Anomaly Detection: Unifying Appearance, Geometry, and Internal Properties
by: Li, Wenqiao, et al.
Published: (2024) -
Bridging 3D Anomaly Localization and Repair via High-Quality Continuous Geometric Representation
by: Zheng, Bozhong, et al.
Published: (2025)