Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929242294452224 |
|---|---|
| author | Salman, Shaeke Shams, Md Montasir Bin Liu, Xiuwen Zhu, Lingjiong |
| author_facet | Salman, Shaeke Shams, Md Montasir Bin Liu, Xiuwen Zhu, Lingjiong |
| contents | Transformer-based models have dominated natural language processing and other areas in the last few years due to their superior (zero-shot) performance on benchmark datasets. However, these models are poorly understood due to their complexity and size. While probing-based methods are widely used to understand specific properties, the structures of the representation space are not systematically characterized; consequently, it is unclear how such models generalize and overgeneralize to new inputs beyond datasets. In this paper, based on a new gradient descent optimization method, we are able to explore the embedding space of a commonly used vision-language model. Using the Imagenette dataset, we show that while the model achieves over 99\% zero-shot classification performance, it fails systematic evaluations completely. Using a linear approximation, we provide a framework to explain the striking differences. We have also obtained similar results using a different model to support that our results are applicable to other transformer models with continuous inputs. We also propose a robust way to detect the modified images. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_08473 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models Salman, Shaeke Shams, Md Montasir Bin Liu, Xiuwen Zhu, Lingjiong Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Transformer-based models have dominated natural language processing and other areas in the last few years due to their superior (zero-shot) performance on benchmark datasets. However, these models are poorly understood due to their complexity and size. While probing-based methods are widely used to understand specific properties, the structures of the representation space are not systematically characterized; consequently, it is unclear how such models generalize and overgeneralize to new inputs beyond datasets. In this paper, based on a new gradient descent optimization method, we are able to explore the embedding space of a commonly used vision-language model. Using the Imagenette dataset, we show that while the model achieves over 99\% zero-shot classification performance, it fails systematic evaluations completely. Using a linear approximation, we provide a framework to explain the striking differences. We have also obtained similar results using a different model to support that our results are applicable to other transformer models with continuous inputs. We also propose a robust way to detect the modified images. |
| title | Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2402.08473 |