Accelerating Vision Foundation Models with Drop-in Depthwise Convolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Scribano, Carmelo, Mahdi, Mohammad, Prisadnikov, Nedyalko, Fu, Yuqian, Franchini, Giorgia, Paudel, Danda Pani, Bertogna, Marko, Van Gool, Luc
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918516283670528
author Scribano, Carmelo
Mahdi, Mohammad
Prisadnikov, Nedyalko
Fu, Yuqian
Franchini, Giorgia
Paudel, Danda Pani
Bertogna, Marko
Van Gool, Luc
author_facet Scribano, Carmelo
Mahdi, Mohammad
Prisadnikov, Nedyalko
Fu, Yuqian
Franchini, Giorgia
Paudel, Danda Pani
Bertogna, Marko
Van Gool, Luc
contents Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones impose high inference costs, limiting deployment on resource-constrained devices. In this work, we accelerate large-scale pretrained ViTs while preserving their feature extraction capabilities by exploiting the intrinsic convolution-like behavior of some attention heads. Specifically, we introduce an efficient depthwise convolution-based layer that serves as a drop-in replacement for these heads. Additionally, we propose simple strategies to identify which heads can be replaced and introduce a fine-tuning procedure that recovers downstream task performance. Across both image classification and segmentation tasks, our method achieves 17-20\% percent inference speedup with minimal performance degradation. We validate the approach through detailed derivations, extensive experiments, and efficiency benchmarks. The reference implementation is publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22132
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
Scribano, Carmelo
Mahdi, Mohammad
Prisadnikov, Nedyalko
Fu, Yuqian
Franchini, Giorgia
Paudel, Danda Pani
Bertogna, Marko
Van Gool, Luc
Computer Vision and Pattern Recognition
Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones impose high inference costs, limiting deployment on resource-constrained devices. In this work, we accelerate large-scale pretrained ViTs while preserving their feature extraction capabilities by exploiting the intrinsic convolution-like behavior of some attention heads. Specifically, we introduce an efficient depthwise convolution-based layer that serves as a drop-in replacement for these heads. Additionally, we propose simple strategies to identify which heads can be replaced and introduce a fine-tuning procedure that recovers downstream task performance. Across both image classification and segmentation tasks, our method achieves 17-20\% percent inference speedup with minimal performance degradation. We validate the approach through detailed derivations, extensive experiments, and efficiency benchmarks. The reference implementation is publicly available.
title Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.22132