Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Jiaqi, Cai, Shaofeng, Deng, Fang, Ooi, Beng Chin, Wu, Junran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024)
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
GAIA: Zero-shot Talking Avatar Generation
von: He, Tianyu, et al.
Veröffentlicht: (2023)
von: He, Tianyu, et al.
Veröffentlicht: (2023)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
METER: A Dynamic Concept Adaptation Framework for Online Anomaly Detection
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2023)
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
von: Liu, Yu, et al.
Veröffentlicht: (2025)
von: Liu, Yu, et al.
Veröffentlicht: (2025)
VAAS: Vision-Attention Anomaly Scoring for Image Manipulation Detection in Digital Forensics
von: Bamigbade, Opeyemi, et al.
Veröffentlicht: (2025)
von: Bamigbade, Opeyemi, et al.
Veröffentlicht: (2025)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
von: Wu, Yike, et al.
Veröffentlicht: (2026)
von: Wu, Yike, et al.
Veröffentlicht: (2026)
Anomaly Detection and Localization for Speech Deepfakes via Feature Pyramid Matching
von: Coletta, Emma, et al.
Veröffentlicht: (2025)
von: Coletta, Emma, et al.
Veröffentlicht: (2025)
Modularized Zero-shot VQA with Pre-trained Models
von: Cao, Rui, et al.
Veröffentlicht: (2023)
von: Cao, Rui, et al.
Veröffentlicht: (2023)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
Selective Vision-Language Subspace Projection for Few-shot CLIP
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
P-GSVC: Layered Progressive 2D Gaussian Splatting for Scalable Image and Video
von: Wang, Longan, et al.
Veröffentlicht: (2026)
von: Wang, Longan, et al.
Veröffentlicht: (2026)
GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting
von: Wang, Longan, et al.
Veröffentlicht: (2025)
von: Wang, Longan, et al.
Veröffentlicht: (2025)
Zero-shot image privacy classification with Vision-Language Models
von: Baia, Alina Elena, et al.
Veröffentlicht: (2025)
von: Baia, Alina Elena, et al.
Veröffentlicht: (2025)
Multimodal Real-Time Anomaly Detection and Industrial Applications
von: Verma, Aman, et al.
Veröffentlicht: (2025)
von: Verma, Aman, et al.
Veröffentlicht: (2025)
Learning Gaussian Data Augmentation in Feature Space for One-shot Object Detection in Manga
von: Taniguchi, Takara, et al.
Veröffentlicht: (2024)
von: Taniguchi, Takara, et al.
Veröffentlicht: (2024)
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2022)
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2022)
Do Joint Audio-Video Generation Models Understand Physics?
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
Nutrition Estimation for Dietary Management: A Transformer Approach with Depth Sensing
von: Kwan, Zhengyi, et al.
Veröffentlicht: (2024)
von: Kwan, Zhengyi, et al.
Veröffentlicht: (2024)
LapisGS: Layered Progressive 3D Gaussian Splatting for Adaptive Streaming
von: Shi, Yuang, et al.
Veröffentlicht: (2024)
von: Shi, Yuang, et al.
Veröffentlicht: (2024)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
von: Yang, Chenglin, et al.
Veröffentlicht: (2023)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
von: Huang, Yiheng, et al.
Veröffentlicht: (2024)
von: Huang, Yiheng, et al.
Veröffentlicht: (2024)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
In Anticipation of Perfect Deepfake: Identity-anchored Artifact-agnostic Detection under Rebalanced Deepfake Detection Protocol
von: Wang, Wei-Han, et al.
Veröffentlicht: (2024)
von: Wang, Wei-Han, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025) -
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024) -
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026) -
GAIA: Zero-shot Talking Avatar Generation
von: He, Tianyu, et al.
Veröffentlicht: (2023) -
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)