Occlusion Robustness of CLIP for Military Vehicle Classification

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: van Woerden, Jan Erik, Burghouts, Gertjan, Nijskens, Lotte, Liezenga, Alma M., van Rooij, Sabina, Ruis, Frank, Kuijf, Hugo J.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912565742796800
author van Woerden, Jan Erik
Burghouts, Gertjan
Nijskens, Lotte
Liezenga, Alma M.
van Rooij, Sabina
Ruis, Frank
Kuijf, Hugo J.
author_facet van Woerden, Jan Erik
Burghouts, Gertjan
Nijskens, Lotte
Liezenga, Alma M.
van Rooij, Sabina
Ruis, Frank
Kuijf, Hugo J.
contents Vision-language models (VLMs) like CLIP enable zero-shot classification by aligning images and text in a shared embedding space, offering advantages for defense applications with scarce labeled data. However, CLIP's robustness in challenging military environments, with partial occlusion and degraded signal-to-noise ratio (SNR), remains underexplored. We investigate CLIP variants' robustness to occlusion using a custom dataset of 18 military vehicle classes and evaluate using Normalized Area Under the Curve (NAUC) across occlusion percentages. Four key insights emerge: (1) Transformer-based CLIP models consistently outperform CNNs, (2) fine-grained, dispersed occlusions degrade performance more than larger contiguous occlusions, (3) despite improved accuracy, performance of linear-probed models sharply drops at around 35% occlusion, (4) by finetuning the model's backbone, this performance drop occurs at more than 60% occlusion. These results underscore the importance of occlusion-specific augmentations during training and the need for further exploration into patch-level sensitivity and architectural resilience for real-world deployment of CLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20760
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Occlusion Robustness of CLIP for Military Vehicle Classification
van Woerden, Jan Erik
Burghouts, Gertjan
Nijskens, Lotte
Liezenga, Alma M.
van Rooij, Sabina
Ruis, Frank
Kuijf, Hugo J.
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-language models (VLMs) like CLIP enable zero-shot classification by aligning images and text in a shared embedding space, offering advantages for defense applications with scarce labeled data. However, CLIP's robustness in challenging military environments, with partial occlusion and degraded signal-to-noise ratio (SNR), remains underexplored. We investigate CLIP variants' robustness to occlusion using a custom dataset of 18 military vehicle classes and evaluate using Normalized Area Under the Curve (NAUC) across occlusion percentages. Four key insights emerge: (1) Transformer-based CLIP models consistently outperform CNNs, (2) fine-grained, dispersed occlusions degrade performance more than larger contiguous occlusions, (3) despite improved accuracy, performance of linear-probed models sharply drops at around 35% occlusion, (4) by finetuning the model's backbone, this performance drop occurs at more than 60% occlusion. These results underscore the importance of occlusion-specific augmentations during training and the need for further exploration into patch-level sensitivity and architectural resilience for real-world deployment of CLIP.
title Occlusion Robustness of CLIP for Military Vehicle Classification
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.20760