EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Xinyu, Sun, Qingyun, Luo, Jiayi, Li, Jianxin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
por: Cheung, Tsun-Hin, et al.
Publicado: (2024)
por: Cheung, Tsun-Hin, et al.
Publicado: (2024)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
por: Zhao, Shuai, et al.
Publicado: (2023)
por: Zhao, Shuai, et al.
Publicado: (2023)
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
por: Duan, Yiqun, et al.
Publicado: (2025)
por: Duan, Yiqun, et al.
Publicado: (2025)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
por: Zhu, Jiaqi, et al.
Publicado: (2024)
por: Zhu, Jiaqi, et al.
Publicado: (2024)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
por: Song, Jiale, et al.
Publicado: (2026)
por: Song, Jiale, et al.
Publicado: (2026)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
por: Cai, Lingling, et al.
Publicado: (2024)
por: Cai, Lingling, et al.
Publicado: (2024)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
por: Wen, Haokun, et al.
Publicado: (2026)
por: Wen, Haokun, et al.
Publicado: (2026)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
por: Fang, Xinyu, et al.
Publicado: (2024)
por: Fang, Xinyu, et al.
Publicado: (2024)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
por: Liao, Liwei, et al.
Publicado: (2025)
por: Liao, Liwei, et al.
Publicado: (2025)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
por: Li, Yingxuan, et al.
Publicado: (2024)
por: Li, Yingxuan, et al.
Publicado: (2024)
Context Guided Transformer Entropy Modeling for Video Compression
por: Tong, Junlong, et al.
Publicado: (2025)
por: Tong, Junlong, et al.
Publicado: (2025)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
por: Lin, Haoqiang, et al.
Publicado: (2025)
por: Lin, Haoqiang, et al.
Publicado: (2025)
Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models
por: Liu, Yang, et al.
Publicado: (2023)
por: Liu, Yang, et al.
Publicado: (2023)
Prompt-Guided Generation of Structured Chest X-Ray Report Using a Pre-trained LLM
por: Li, Hongzhao, et al.
Publicado: (2024)
por: Li, Hongzhao, et al.
Publicado: (2024)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
por: Tan, Jiawei, et al.
Publicado: (2024)
por: Tan, Jiawei, et al.
Publicado: (2024)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
por: He, Jiayi, et al.
Publicado: (2025)
por: He, Jiayi, et al.
Publicado: (2025)
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
por: Yang, Danni, et al.
Publicado: (2024)
por: Yang, Danni, et al.
Publicado: (2024)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
por: Shi, Haoyuan, et al.
Publicado: (2026)
por: Shi, Haoyuan, et al.
Publicado: (2026)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
por: Hsiao, Teng-Fang, et al.
Publicado: (2024)
por: Hsiao, Teng-Fang, et al.
Publicado: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
por: Miao, Bingchen, et al.
Publicado: (2024)
por: Miao, Bingchen, et al.
Publicado: (2024)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
por: Gao, Jiayi, et al.
Publicado: (2025)
por: Gao, Jiayi, et al.
Publicado: (2025)
LAPIG: Language Guided Projector Image Generation with Surface Adaptation and Stylization
por: Deng, Yuchen, et al.
Publicado: (2025)
por: Deng, Yuchen, et al.
Publicado: (2025)
DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera Synthesis
por: Wang, Zixuan, et al.
Publicado: (2024)
por: Wang, Zixuan, et al.
Publicado: (2024)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
por: Xie, Zequn, et al.
Publicado: (2026)
por: Xie, Zequn, et al.
Publicado: (2026)
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
por: Liu, Yu, et al.
Publicado: (2025)
por: Liu, Yu, et al.
Publicado: (2025)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
por: Cui, Can, et al.
Publicado: (2024)
por: Cui, Can, et al.
Publicado: (2024)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
por: Hoe, Jiun Tian, et al.
Publicado: (2025)
por: Hoe, Jiun Tian, et al.
Publicado: (2025)
Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
por: Zhu, Sa, et al.
Publicado: (2026)
por: Zhu, Sa, et al.
Publicado: (2026)
VAAS: Vision-Attention Anomaly Scoring for Image Manipulation Detection in Digital Forensics
por: Bamigbade, Opeyemi, et al.
Publicado: (2025)
por: Bamigbade, Opeyemi, et al.
Publicado: (2025)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
por: Dai, Guangyu, et al.
Publicado: (2025)
por: Dai, Guangyu, et al.
Publicado: (2025)
Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
por: Cai, Haonan, et al.
Publicado: (2026)
por: Cai, Haonan, et al.
Publicado: (2026)
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
por: Li, Shaokai, et al.
Publicado: (2024)
por: Li, Shaokai, et al.
Publicado: (2024)
Visual Grounding with Multi-modal Conditional Adaptation
por: Yao, Ruilin, et al.
Publicado: (2024)
por: Yao, Ruilin, et al.
Publicado: (2024)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
por: Luo, Jianjie, et al.
Publicado: (2024)
por: Luo, Jianjie, et al.
Publicado: (2024)
Towards Generalizable Deepfake Detection via Forgery-aware Audio-Visual Adaptation: A Variational Bayesian Approach
por: Nie, Fan, et al.
Publicado: (2025)
por: Nie, Fan, et al.
Publicado: (2025)
Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling
por: Li, Xinyu, et al.
Publicado: (2025)
por: Li, Xinyu, et al.
Publicado: (2025)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
por: Xu, Yifang, et al.
Publicado: (2025)
por: Xu, Yifang, et al.
Publicado: (2025)
An Effective Image Copy-Move Forgery Detection Using Entropy Information
por: Jiang, Li, et al.
Publicado: (2023)
por: Jiang, Li, et al.
Publicado: (2023)
GAIA: Zero-shot Talking Avatar Generation
por: He, Tianyu, et al.
Publicado: (2023)
por: He, Tianyu, et al.
Publicado: (2023)
KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection
por: Li, Xingyuan, et al.
Publicado: (2025)
por: Li, Xingyuan, et al.
Publicado: (2025)
Ejemplares similares
-
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
por: Cheung, Tsun-Hin, et al.
Publicado: (2024) -
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
por: Zhao, Shuai, et al.
Publicado: (2023) -
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
por: Duan, Yiqun, et al.
Publicado: (2025) -
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
por: Zhu, Jiaqi, et al.
Publicado: (2024) -
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
por: Song, Jiale, et al.
Publicado: (2026)