SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Shu, Tian, Xinyu, Zhao, Qinyu, Yang, Zhaoyuan, Zhang, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909462661431296
author Zou, Shu
Tian, Xinyu
Zhao, Qinyu
Yang, Zhaoyuan
Zhang, Jing
author_facet Zou, Shu
Tian, Xinyu
Zhao, Qinyu
Yang, Zhaoyuan
Zhang, Jing
contents Detecting out-of-distribution (OOD) data is crucial in real-world machine learning applications, particularly in safety-critical domains. Existing methods often leverage language information from vision-language models (VLMs) to enhance OOD detection by improving confidence estimation through rich class-wise text information. However, when building OOD detection score upon on in-distribution (ID) text-image affinity, existing works either focus on each ID class or whole ID label sets, overlooking inherent ID classes' connection. We find that the semantic information across different ID classes is beneficial for effective OOD detection. We thus investigate the ability of image-text comprehension among different semantic-related ID labels in VLMs and propose a novel post-hoc strategy called SimLabel. SimLabel enhances the separability between ID and OOD samples by establishing a more robust image-class similarity metric that considers consistency over a set of similar class labels. Extensive experiments demonstrate the superior performance of SimLabel on various zero-shot OOD detection benchmarks. The proposed model is also extended to various VLM-backbones, demonstrating its good generalization ability. Our demonstration and implementation codes are available at: https://github.com/ShuZou-1/SimLabel.
format Preprint
id arxiv_https___arxiv_org_abs_2501_11485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models
Zou, Shu
Tian, Xinyu
Zhao, Qinyu
Yang, Zhaoyuan
Zhang, Jing
Computer Vision and Pattern Recognition
Detecting out-of-distribution (OOD) data is crucial in real-world machine learning applications, particularly in safety-critical domains. Existing methods often leverage language information from vision-language models (VLMs) to enhance OOD detection by improving confidence estimation through rich class-wise text information. However, when building OOD detection score upon on in-distribution (ID) text-image affinity, existing works either focus on each ID class or whole ID label sets, overlooking inherent ID classes' connection. We find that the semantic information across different ID classes is beneficial for effective OOD detection. We thus investigate the ability of image-text comprehension among different semantic-related ID labels in VLMs and propose a novel post-hoc strategy called SimLabel. SimLabel enhances the separability between ID and OOD samples by establishing a more robust image-class similarity metric that considers consistency over a set of similar class labels. Extensive experiments demonstrate the superior performance of SimLabel on various zero-shot OOD detection benchmarks. The proposed model is also extended to various VLM-backbones, demonstrating its good generalization ability. Our demonstration and implementation codes are available at: https://github.com/ShuZou-1/SimLabel.
title SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.11485