Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911750783238144 |
|---|---|
| author | Alabdulmohsin, Ibrahim Zhai, Xiaohua Kolesnikov, Alexander Beyer, Lucas |
| author_facet | Alabdulmohsin, Ibrahim Zhai, Xiaohua Kolesnikov, Alexander Beyer, Lucas |
| contents | Scaling laws have been recently employed to derive compute-optimal model size (number of parameters) for a given compute duration. We advance and refine such methods to infer compute-optimal model shapes, such as width and depth, and successfully implement this in vision transformers. Our shape-optimized vision transformer, SoViT, achieves results competitive with models that exceed twice its size, despite being pre-trained with an equivalent amount of compute. For example, SoViT-400m/14 achieves 90.3% fine-tuning accuracy on ILSRCV2012, surpassing the much larger ViT-g/14 and approaching ViT-G/14 under identical settings, with also less than half the inference cost. We conduct a thorough evaluation across multiple tasks, such as image classification, captioning, VQA and zero-shot transfer, demonstrating the effectiveness of our model across a broad range of domains and identifying limitations. Overall, our findings challenge the prevailing approach of blindly scaling up vision models and pave a path for a more informed scaling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_13035 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design Alabdulmohsin, Ibrahim Zhai, Xiaohua Kolesnikov, Alexander Beyer, Lucas Computer Vision and Pattern Recognition Machine Learning I.2.10; I.2.6 Scaling laws have been recently employed to derive compute-optimal model size (number of parameters) for a given compute duration. We advance and refine such methods to infer compute-optimal model shapes, such as width and depth, and successfully implement this in vision transformers. Our shape-optimized vision transformer, SoViT, achieves results competitive with models that exceed twice its size, despite being pre-trained with an equivalent amount of compute. For example, SoViT-400m/14 achieves 90.3% fine-tuning accuracy on ILSRCV2012, surpassing the much larger ViT-g/14 and approaching ViT-G/14 under identical settings, with also less than half the inference cost. We conduct a thorough evaluation across multiple tasks, such as image classification, captioning, VQA and zero-shot transfer, demonstrating the effectiveness of our model across a broad range of domains and identifying limitations. Overall, our findings challenge the prevailing approach of blindly scaling up vision models and pave a path for a more informed scaling. |
| title | Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design |
| topic | Computer Vision and Pattern Recognition Machine Learning I.2.10; I.2.6 |
| url | https://arxiv.org/abs/2305.13035 |