Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Shizhen, Liu, Jiahui, Wen, Xin, Tan, Haoru, Qi, Xiaojuan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914089563848704
author Zhao, Shizhen
Liu, Jiahui
Wen, Xin
Tan, Haoru
Qi, Xiaojuan
author_facet Zhao, Shizhen
Liu, Jiahui
Wen, Xin
Tan, Haoru
Qi, Xiaojuan
contents Pre-trained vision foundation models have transformed many computer vision tasks. Despite their strong ability to learn discriminative and generalizable features crucial for out-of-distribution (OOD) detection, their impact on this task remains underexplored. Motivated by this gap, we systematically investigate representative vision foundation models for OOD detection. Our findings reveal that a pre-trained DINOv2 model, even without fine-tuning on in-domain (ID) data, naturally provides a highly discriminative feature space for OOD detection, achieving performance comparable to existing state-of-the-art methods without requiring complex designs. Beyond this, we explore how fine-tuning foundation models on in-domain (ID) data can enhance OOD detection. However, we observe that the performance of vision foundation models remains unsatisfactory in scenarios with a large semantic space. This is due to the increased complexity of decision boundaries as the number of categories grows, which complicates the optimization process. To mitigate this, we propose the Mixture of Feature Experts (MoFE) module, which partitions features into subspaces, effectively capturing complex data distributions and refining decision boundaries. Further, we introduce a Dynamic-$β$ Mixup strategy, which samples interpolation weights from a dynamic beta distribution. This adapts to varying levels of learning difficulty across categories, improving feature learning for more challenging categories. Extensive experiments demonstrate the effectiveness of our approach, significantly outperforming baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
Zhao, Shizhen
Liu, Jiahui
Wen, Xin
Tan, Haoru
Qi, Xiaojuan
Computer Vision and Pattern Recognition
Pre-trained vision foundation models have transformed many computer vision tasks. Despite their strong ability to learn discriminative and generalizable features crucial for out-of-distribution (OOD) detection, their impact on this task remains underexplored. Motivated by this gap, we systematically investigate representative vision foundation models for OOD detection. Our findings reveal that a pre-trained DINOv2 model, even without fine-tuning on in-domain (ID) data, naturally provides a highly discriminative feature space for OOD detection, achieving performance comparable to existing state-of-the-art methods without requiring complex designs. Beyond this, we explore how fine-tuning foundation models on in-domain (ID) data can enhance OOD detection. However, we observe that the performance of vision foundation models remains unsatisfactory in scenarios with a large semantic space. This is due to the increased complexity of decision boundaries as the number of categories grows, which complicates the optimization process. To mitigate this, we propose the Mixture of Feature Experts (MoFE) module, which partitions features into subspaces, effectively capturing complex data distributions and refining decision boundaries. Further, we introduce a Dynamic-$β$ Mixup strategy, which samples interpolation weights from a dynamic beta distribution. This adapts to varying levels of learning difficulty across categories, improving feature learning for more challenging categories. Extensive experiments demonstrate the effectiveness of our approach, significantly outperforming baseline methods.
title Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10584