Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918516283670528 |
|---|---|
| author | Scribano, Carmelo Mahdi, Mohammad Prisadnikov, Nedyalko Fu, Yuqian Franchini, Giorgia Paudel, Danda Pani Bertogna, Marko Van Gool, Luc |
| author_facet | Scribano, Carmelo Mahdi, Mohammad Prisadnikov, Nedyalko Fu, Yuqian Franchini, Giorgia Paudel, Danda Pani Bertogna, Marko Van Gool, Luc |
| contents | Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones impose high inference costs, limiting deployment on resource-constrained devices. In this work, we accelerate large-scale pretrained ViTs while preserving their feature extraction capabilities by exploiting the intrinsic convolution-like behavior of some attention heads. Specifically, we introduce an efficient depthwise convolution-based layer that serves as a drop-in replacement for these heads. Additionally, we propose simple strategies to identify which heads can be replaced and introduce a fine-tuning procedure that recovers downstream task performance. Across both image classification and segmentation tasks, our method achieves 17-20\% percent inference speedup with minimal performance degradation. We validate the approach through detailed derivations, extensive experiments, and efficiency benchmarks. The reference implementation is publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_22132 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Accelerating Vision Foundation Models with Drop-in Depthwise Convolution Scribano, Carmelo Mahdi, Mohammad Prisadnikov, Nedyalko Fu, Yuqian Franchini, Giorgia Paudel, Danda Pani Bertogna, Marko Van Gool, Luc Computer Vision and Pattern Recognition Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones impose high inference costs, limiting deployment on resource-constrained devices. In this work, we accelerate large-scale pretrained ViTs while preserving their feature extraction capabilities by exploiting the intrinsic convolution-like behavior of some attention heads. Specifically, we introduce an efficient depthwise convolution-based layer that serves as a drop-in replacement for these heads. Additionally, we propose simple strategies to identify which heads can be replaced and introduce a fine-tuning procedure that recovers downstream task performance. Across both image classification and segmentation tasks, our method achieves 17-20\% percent inference speedup with minimal performance degradation. We validate the approach through detailed derivations, extensive experiments, and efficiency benchmarks. The reference implementation is publicly available. |
| title | Accelerating Vision Foundation Models with Drop-in Depthwise Convolution |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.22132 |