Benchmarking Feature Upsampling Methods for Vision Foundation Models using Interactive Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Havrylov, Volodymyr, Huang, Haiwen, Zhang, Dan, Geiger, Andreas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908349746905088
author Havrylov, Volodymyr
Huang, Haiwen
Zhang, Dan
Geiger, Andreas
author_facet Havrylov, Volodymyr
Huang, Haiwen
Zhang, Dan
Geiger, Andreas
contents Vision Foundation Models (VFMs) are large-scale, pre-trained models that serve as general-purpose backbones for various computer vision tasks. As VFMs' popularity grows, there is an increasing interest in understanding their effectiveness for dense prediction tasks. However, VFMs typically produce low-resolution features, limiting their direct applicability in this context. One way to tackle this limitation is by employing a task-agnostic feature upsampling module that refines VFM features resolution. To assess the effectiveness of this approach, we investigate Interactive Segmentation (IS) as a novel benchmark for evaluating feature upsampling methods on VFMs. Due to its inherent multimodal input, consisting of an image and a set of user-defined clicks, as well as its dense mask output, IS creates a challenging environment that demands comprehensive visual scene understanding. Our benchmarking experiments show that selecting appropriate upsampling strategies significantly improves VFM features quality. The code is released at https://github.com/havrylovv/iSegProbe
format Preprint
id arxiv_https___arxiv_org_abs_2505_02075
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Feature Upsampling Methods for Vision Foundation Models using Interactive Segmentation
Havrylov, Volodymyr
Huang, Haiwen
Zhang, Dan
Geiger, Andreas
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Vision Foundation Models (VFMs) are large-scale, pre-trained models that serve as general-purpose backbones for various computer vision tasks. As VFMs' popularity grows, there is an increasing interest in understanding their effectiveness for dense prediction tasks. However, VFMs typically produce low-resolution features, limiting their direct applicability in this context. One way to tackle this limitation is by employing a task-agnostic feature upsampling module that refines VFM features resolution. To assess the effectiveness of this approach, we investigate Interactive Segmentation (IS) as a novel benchmark for evaluating feature upsampling methods on VFMs. Due to its inherent multimodal input, consisting of an image and a set of user-defined clicks, as well as its dense mask output, IS creates a challenging environment that demands comprehensive visual scene understanding. Our benchmarking experiments show that selecting appropriate upsampling strategies significantly improves VFM features quality. The code is released at https://github.com/havrylovv/iSegProbe
title Benchmarking Feature Upsampling Methods for Vision Foundation Models using Interactive Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.02075