TIPS Over Tricks: Simple Prompts for Effective Zero-shot Anomaly Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911468231852032 |
|---|---|
| author | Salehi, Alireza Karami, Ehsan Noey, Sepehr Noey, Sahand Yamada, Makoto Hosseini, Reshad Sabokrou, Mohammad |
| author_facet | Salehi, Alireza Karami, Ehsan Noey, Sepehr Noey, Sahand Yamada, Makoto Hosseini, Reshad Sabokrou, Mohammad |
| contents | Anomaly detection identifies departures from expected behavior in safety-critical settings. When target-domain normal data are unavailable, zero-shot anomaly detection (ZSAD) leverages vision-language models (VLMs). However, CLIP's coarse image-text alignment limits both localization and detection due to (i) spatial misalignment and (ii) weak sensitivity to fine-grained anomalies; prior works compensate with complex auxiliary modules yet largely overlook the choice of backbone. We revisit the backbone and use TIPS-a VLM trained with spatially aware objectives. While TIPS alleviates CLIP's issues, it exposes a distributional gap between global and local features. We address this with decoupled prompts-fixed for image-level detection and learnable for pixel-level localization-and by injecting local evidence into the global score. Without CLIP-specific tricks, our TIPS-based pipeline improves image-level performance by 1.1-3.9% and pixel-level by 1.5-6.9% across seven industrial datasets, delivering strong generalization with a lean architecture. Code is available at github.com/AlirezaSalehy/Tipsomaly. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_03594 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | TIPS Over Tricks: Simple Prompts for Effective Zero-shot Anomaly Detection Salehi, Alireza Karami, Ehsan Noey, Sepehr Noey, Sahand Yamada, Makoto Hosseini, Reshad Sabokrou, Mohammad Computer Vision and Pattern Recognition Anomaly detection identifies departures from expected behavior in safety-critical settings. When target-domain normal data are unavailable, zero-shot anomaly detection (ZSAD) leverages vision-language models (VLMs). However, CLIP's coarse image-text alignment limits both localization and detection due to (i) spatial misalignment and (ii) weak sensitivity to fine-grained anomalies; prior works compensate with complex auxiliary modules yet largely overlook the choice of backbone. We revisit the backbone and use TIPS-a VLM trained with spatially aware objectives. While TIPS alleviates CLIP's issues, it exposes a distributional gap between global and local features. We address this with decoupled prompts-fixed for image-level detection and learnable for pixel-level localization-and by injecting local evidence into the global score. Without CLIP-specific tricks, our TIPS-based pipeline improves image-level performance by 1.1-3.9% and pixel-level by 1.5-6.9% across seven industrial datasets, delivering strong generalization with a lean architecture. Code is available at github.com/AlirezaSalehy/Tipsomaly. |
| title | TIPS Over Tricks: Simple Prompts for Effective Zero-shot Anomaly Detection |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.03594 |