Anomaly Detection by Adapting a pre-trained Vision Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Yuxuan, He, Xinwei, Liang, Dingkang, Tong, Ao, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omni-AD: Learning to Reconstruct Global and Local Features for Multi-class Anomaly Detection
by: Quan, Jiajie, et al.
Published: (2025)
by: Quan, Jiajie, et al.
Published: (2025)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
by: He, Xinwei, et al.
Published: (2026)
by: He, Xinwei, et al.
Published: (2026)
Not All Regions Are Equal: Attention-Guided Perturbation Network for Industrial Anomaly Detection
by: Huang, Tingfeng, et al.
Published: (2024)
by: Huang, Tingfeng, et al.
Published: (2024)
Adapting Vision-Language Models to Open Classes via Test-Time Prompt Tuning
by: Gao, Zhengqing, et al.
Published: (2024)
by: Gao, Zhengqing, et al.
Published: (2024)
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
by: Zhao, Zongchuang, et al.
Published: (2025)
by: Zhao, Zongchuang, et al.
Published: (2025)
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
by: Lin, Hongkai, et al.
Published: (2025)
by: Lin, Hongkai, et al.
Published: (2025)
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Adapting Visual-Language Models for Generalizable Anomaly Detection in Medical Images
by: Huang, Chaoqin, et al.
Published: (2024)
by: Huang, Chaoqin, et al.
Published: (2024)
Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
by: Wang, Zhichuan, et al.
Published: (2025)
by: Wang, Zhichuan, et al.
Published: (2025)
More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
by: Lin, Hongkai, et al.
Published: (2025)
by: Lin, Hongkai, et al.
Published: (2025)
SOOD++: Leveraging Unlabeled Data to Boost Oriented Object Detection
by: Liang, Dingkang, et al.
Published: (2024)
by: Liang, Dingkang, et al.
Published: (2024)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
by: Lei, Yuxuan, et al.
Published: (2024)
by: Lei, Yuxuan, et al.
Published: (2024)
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
EventVAD: Training-Free Event-Aware Video Anomaly Detection
by: Shao, Yihua, et al.
Published: (2025)
by: Shao, Yihua, et al.
Published: (2025)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection
by: Khan, Sohail Ahmed, et al.
Published: (2024)
by: Khan, Sohail Ahmed, et al.
Published: (2024)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
by: Zhu, Xingkui, et al.
Published: (2024)
by: Zhu, Xingkui, et al.
Published: (2024)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
Vision-Language Models Assisted Unsupervised Video Anomaly Detection
by: Jiang, Yalong, et al.
Published: (2024)
by: Jiang, Yalong, et al.
Published: (2024)
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
by: Lu, Yangxiao, et al.
Published: (2024)
by: Lu, Yangxiao, et al.
Published: (2024)
Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis
by: Qu, Jingguo, et al.
Published: (2025)
by: Qu, Jingguo, et al.
Published: (2025)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2026)
by: Zhou, Xin, et al.
Published: (2026)
SOWA: Adapting Hierarchical Frozen Window Self-Attention to Visual-Language Models for Better Anomaly Detection
by: Hu, Zongxiang, et al.
Published: (2024)
by: Hu, Zongxiang, et al.
Published: (2024)
MAA: Meticulous Adversarial Attack against Vision-Language Pre-trained Models
by: Zhang, Peng-Fei, et al.
Published: (2025)
by: Zhang, Peng-Fei, et al.
Published: (2025)
Towards Training-free Anomaly Detection with Vision and Language Foundation Models
by: Zhang, Jinjin, et al.
Published: (2025)
by: Zhang, Jinjin, et al.
Published: (2025)
NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
by: Xu, Wei, et al.
Published: (2025)
by: Xu, Wei, et al.
Published: (2025)
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
by: Tu, Sifan, et al.
Published: (2025)
by: Tu, Sifan, et al.
Published: (2025)
You Only Look Bottom-Up for Monocular 3D Object Detection
by: Xiong, Kaixin, et al.
Published: (2024)
by: Xiong, Kaixin, et al.
Published: (2024)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
by: Sun, Zhengyang, et al.
Published: (2026)
by: Sun, Zhengyang, et al.
Published: (2026)
TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment
by: Wang, Zhichuan, et al.
Published: (2025)
by: Wang, Zhichuan, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning
by: Liang, Dingkang, et al.
Published: (2024)
by: Liang, Dingkang, et al.
Published: (2024)
A Unified Framework for 3D Scene Understanding
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
MINIMA: Modality Invariant Image Matching
by: Ren, Jiangwei, et al.
Published: (2024)
by: Ren, Jiangwei, et al.
Published: (2024)
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
by: Liang, Dingkang, et al.
Published: (2025)
by: Liang, Dingkang, et al.
Published: (2025)
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection
by: Qian, Kun, et al.
Published: (2024)
by: Qian, Kun, et al.
Published: (2024)
Similar Items
-
Omni-AD: Learning to Reconstruct Global and Local Features for Multi-class Anomaly Detection
by: Quan, Jiajie, et al.
Published: (2025) -
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
by: Cai, Yuxuan, et al.
Published: (2024) -
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
by: He, Xinwei, et al.
Published: (2026) -
Not All Regions Are Equal: Attention-Guided Perturbation Network for Industrial Anomaly Detection
by: Huang, Tingfeng, et al.
Published: (2024) -
Adapting Vision-Language Models to Open Classes via Test-Time Prompt Tuning
by: Gao, Zhengqing, et al.
Published: (2024)