FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Stephanie, Hamilton, Mark, Brandt, Laura, Feldman, Axel, Zhang, Zhoutong, Freeman, William T.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910394886389760
author Fu, Stephanie
Hamilton, Mark
Brandt, Laura
Feldman, Axel
Zhang, Zhoutong
Freeman, William T.
author_facet Fu, Stephanie
Hamilton, Mark
Brandt, Laura
Feldman, Axel
Zhang, Zhoutong
Freeman, William T.
contents Deep features are a cornerstone of computer vision research, capturing image semantics and enabling the community to solve downstream tasks even in the zero- or few-shot regime. However, these features often lack the spatial resolution to directly perform dense prediction tasks like segmentation and depth prediction because models aggressively pool information over large areas. In this work, we introduce FeatUp, a task- and model-agnostic framework to restore lost spatial information in deep features. We introduce two variants of FeatUp: one that guides features with high-resolution signal in a single forward pass, and one that fits an implicit model to a single image to reconstruct features at any resolution. Both approaches use a multi-view consistency loss with deep analogies to NeRFs. Our features retain their original semantics and can be swapped into existing applications to yield resolution and performance gains even without re-training. We show that FeatUp significantly outperforms other feature upsampling and image super-resolution approaches in class activation map generation, transfer learning for segmentation and depth prediction, and end-to-end training for semantic segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10516
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FeatUp: A Model-Agnostic Framework for Features at Any Resolution
Fu, Stephanie
Hamilton, Mark
Brandt, Laura
Feldman, Axel
Zhang, Zhoutong
Freeman, William T.
Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
Machine Learning
Deep features are a cornerstone of computer vision research, capturing image semantics and enabling the community to solve downstream tasks even in the zero- or few-shot regime. However, these features often lack the spatial resolution to directly perform dense prediction tasks like segmentation and depth prediction because models aggressively pool information over large areas. In this work, we introduce FeatUp, a task- and model-agnostic framework to restore lost spatial information in deep features. We introduce two variants of FeatUp: one that guides features with high-resolution signal in a single forward pass, and one that fits an implicit model to a single image to reconstruct features at any resolution. Both approaches use a multi-view consistency loss with deep analogies to NeRFs. Our features retain their original semantics and can be swapped into existing applications to yield resolution and performance gains even without re-training. We show that FeatUp significantly outperforms other feature upsampling and image super-resolution approaches in class activation map generation, transfer learning for segmentation and depth prediction, and end-to-end training for semantic segmentation.
title FeatUp: A Model-Agnostic Framework for Features at Any Resolution
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2403.10516