Weighted Reverse Convolution for Feature Upsampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wentong, Qi, Zhiyuan, Zhao, Zichen, Zhang, Kai, Zhang, Lei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918512822321152
author Li, Wentong
Qi, Zhiyuan
Zhao, Zichen
Zhang, Kai
Zhang, Lei
author_facet Li, Wentong
Qi, Zhiyuan
Zhao, Zichen
Zhang, Kai
Zhang, Lei
contents Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting their effectiveness on tasks requiring fine-grained localization, dense prediction, and point-wise correspondence. In this work, we revisit feature upsampling for VFMs from the perspective of \textbf{\textit{inverse problem}} and propose Weighted Reverse Convolution (WRC), a spatially adaptive inverse operator for densifying high-level visual descriptors. Specifically, we formulate feature upsampling as a weighted Tikhonov-regularized least-squares problem, where spatially varying weights modulate both data fidelity and prior strength at each spatial location. This allows WRC to adapt the reconstruction to spatially varying feature characteristics, thereby preserving critical structures while mitigating over-smoothing. Moreover, WRC retains an efficient, fully differentiable closed-form FFT solution, making it a practical drop-in upsampling operator. Integrated into a lightweight self-supervised densification framework, WRC consistently improves dense feature quality across various downstream benchmarks, including segmentation, depth estimation, video object segmentation, object discovery, and keypoint correspondence, while maintaining high computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17472
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Weighted Reverse Convolution for Feature Upsampling
Li, Wentong
Qi, Zhiyuan
Zhao, Zichen
Zhang, Kai
Zhang, Lei
Computer Vision and Pattern Recognition
Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting their effectiveness on tasks requiring fine-grained localization, dense prediction, and point-wise correspondence. In this work, we revisit feature upsampling for VFMs from the perspective of \textbf{\textit{inverse problem}} and propose Weighted Reverse Convolution (WRC), a spatially adaptive inverse operator for densifying high-level visual descriptors. Specifically, we formulate feature upsampling as a weighted Tikhonov-regularized least-squares problem, where spatially varying weights modulate both data fidelity and prior strength at each spatial location. This allows WRC to adapt the reconstruction to spatially varying feature characteristics, thereby preserving critical structures while mitigating over-smoothing. Moreover, WRC retains an efficient, fully differentiable closed-form FFT solution, making it a practical drop-in upsampling operator. Integrated into a lightweight self-supervised densification framework, WRC consistently improves dense feature quality across various downstream benchmarks, including segmentation, depth estimation, video object segmentation, object discovery, and keypoint correspondence, while maintaining high computational efficiency.
title Weighted Reverse Convolution for Feature Upsampling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17472