NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chambon, Loick, Couairon, Paul, Zablocki, Eloi, Boulch, Alexandre, Thome, Nicolas, Cord, Matthieu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917100005621760
author Chambon, Loick
Couairon, Paul
Zablocki, Eloi
Boulch, Alexandre
Thome, Nicolas
Cord, Matthieu
author_facet Chambon, Loick
Couairon, Paul
Zablocki, Eloi
Boulch, Alexandre
Thome, Nicolas
Cord, Matthieu
contents Vision Foundation Models (VFMs) extract spatially downsampled representations, posing challenges for pixel-level tasks. Existing upsampling approaches face a fundamental trade-off: classical filters are fast and broadly applicable but rely on fixed forms, while modern upsamplers achieve superior accuracy through learnable, VFM-specific forms at the cost of retraining for each VFM. We introduce Neighborhood Attention Filtering (NAF), which bridges this gap by learning adaptive spatial-and-content weights through Cross-Scale Neighborhood Attention and Rotary Position Embeddings (RoPE), guided solely by the high-resolution input image. NAF operates zero-shot: it upsamples features from any VFM without retraining, making it the first VFM-agnostic architecture to outperform VFM-specific upsamplers and achieve state-of-the-art performance across multiple downstream tasks. It maintains high efficiency, scaling to 2K feature maps and reconstructing intermediate-resolution maps at 18 FPS. Beyond feature upsampling, NAF demonstrates strong performance on image restoration, highlighting its versatility. Code and checkpoints are available at https://github.com/valeoai/NAF.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18452
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
Chambon, Loick
Couairon, Paul
Zablocki, Eloi
Boulch, Alexandre
Thome, Nicolas
Cord, Matthieu
Computer Vision and Pattern Recognition
Vision Foundation Models (VFMs) extract spatially downsampled representations, posing challenges for pixel-level tasks. Existing upsampling approaches face a fundamental trade-off: classical filters are fast and broadly applicable but rely on fixed forms, while modern upsamplers achieve superior accuracy through learnable, VFM-specific forms at the cost of retraining for each VFM. We introduce Neighborhood Attention Filtering (NAF), which bridges this gap by learning adaptive spatial-and-content weights through Cross-Scale Neighborhood Attention and Rotary Position Embeddings (RoPE), guided solely by the high-resolution input image. NAF operates zero-shot: it upsamples features from any VFM without retraining, making it the first VFM-agnostic architecture to outperform VFM-specific upsamplers and achieve state-of-the-art performance across multiple downstream tasks. It maintains high efficiency, scaling to 2K feature maps and reconstructing intermediate-resolution maps at 18 FPS. Beyond feature upsampling, NAF demonstrates strong performance on image restoration, highlighting its versatility. Code and checkpoints are available at https://github.com/valeoai/NAF.
title NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18452