FeatSharp: Your Vision Model Features, Sharper

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ranzinger, Mike, Heinrich, Greg, Molchanov, Pavlo, Kautz, Jan, Catanzaro, Bryan, Tao, Andrew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916823455236096
author Ranzinger, Mike
Heinrich, Greg
Molchanov, Pavlo
Kautz, Jan
Catanzaro, Bryan
Tao, Andrew
author_facet Ranzinger, Mike
Heinrich, Greg
Molchanov, Pavlo
Kautz, Jan
Catanzaro, Bryan
Tao, Andrew
contents The feature maps of vision encoders are fundamental to myriad modern AI tasks, ranging from core perception algorithms (e.g. semantic segmentation, object detection, depth perception, etc.) to modern multimodal understanding in vision-language models (VLMs). Currently, in computer vision, the frontier of general purpose vision backbones is Vision Transformers (ViT), typically trained using contrastive loss (e.g. CLIP). A key problem with most off-the-shelf ViTs, particularly CLIP, is that these models are inflexibly low resolution. Most run at $224 \times 224$px, while the "high-resolution" versions are around $378-448$px, but still inflexible. We introduce a novel method to coherently and cheaply upsample the feature maps of low-resolution vision encoders while picking up on fine-grained details that would otherwise be lost due to resolution. We demonstrate the effectiveness of this approach on core perception tasks as well as within agglomerative model training using RADIO as a way of providing richer targets for distillation. Code available at https://github.com/NVlabs/FeatSharp .
format Preprint
id arxiv_https___arxiv_org_abs_2502_16025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FeatSharp: Your Vision Model Features, Sharper
Ranzinger, Mike
Heinrich, Greg
Molchanov, Pavlo
Kautz, Jan
Catanzaro, Bryan
Tao, Andrew
Computer Vision and Pattern Recognition
The feature maps of vision encoders are fundamental to myriad modern AI tasks, ranging from core perception algorithms (e.g. semantic segmentation, object detection, depth perception, etc.) to modern multimodal understanding in vision-language models (VLMs). Currently, in computer vision, the frontier of general purpose vision backbones is Vision Transformers (ViT), typically trained using contrastive loss (e.g. CLIP). A key problem with most off-the-shelf ViTs, particularly CLIP, is that these models are inflexibly low resolution. Most run at $224 \times 224$px, while the "high-resolution" versions are around $378-448$px, but still inflexible. We introduce a novel method to coherently and cheaply upsample the feature maps of low-resolution vision encoders while picking up on fine-grained details that would otherwise be lost due to resolution. We demonstrate the effectiveness of this approach on core perception tasks as well as within agglomerative model training using RADIO as a way of providing richer targets for distillation. Code available at https://github.com/NVlabs/FeatSharp .
title FeatSharp: Your Vision Model Features, Sharper
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.16025