Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Grace, Dunlap, Lisa, Park, Dong Huk, Holynski, Aleksander, Darrell, Trevor
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917627774894080
author Luo, Grace
Dunlap, Lisa
Park, Dong Huk
Holynski, Aleksander
Darrell, Trevor
author_facet Luo, Grace
Dunlap, Lisa
Park, Dong Huk
Holynski, Aleksander
Darrell, Trevor
contents Diffusion models have been shown to be capable of generating high-quality images, suggesting that they could contain meaningful internal representations. Unfortunately, the feature maps that encode a diffusion model's internal information are spread not only over layers of the network, but also over diffusion timesteps, making it challenging to extract useful descriptors. We propose Diffusion Hyperfeatures, a framework for consolidating multi-scale and multi-timestep feature maps into per-pixel feature descriptors that can be used for downstream tasks. These descriptors can be extracted for both synthetic and real images using the generation and inversion processes. We evaluate the utility of our Diffusion Hyperfeatures on the task of semantic keypoint correspondence: our method achieves superior performance on the SPair-71k real image benchmark. We also demonstrate that our method is flexible and transferable: our feature aggregation network trained on the inversion features of real image pairs can be used on the generation features of synthetic image pairs with unseen objects and compositions. Our code is available at https://diffusion-hyperfeatures.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14334
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
Luo, Grace
Dunlap, Lisa
Park, Dong Huk
Holynski, Aleksander
Darrell, Trevor
Computer Vision and Pattern Recognition
Diffusion models have been shown to be capable of generating high-quality images, suggesting that they could contain meaningful internal representations. Unfortunately, the feature maps that encode a diffusion model's internal information are spread not only over layers of the network, but also over diffusion timesteps, making it challenging to extract useful descriptors. We propose Diffusion Hyperfeatures, a framework for consolidating multi-scale and multi-timestep feature maps into per-pixel feature descriptors that can be used for downstream tasks. These descriptors can be extracted for both synthetic and real images using the generation and inversion processes. We evaluate the utility of our Diffusion Hyperfeatures on the task of semantic keypoint correspondence: our method achieves superior performance on the SPair-71k real image benchmark. We also demonstrate that our method is flexible and transferable: our feature aggregation network trained on the inversion features of real image pairs can be used on the generation features of synthetic image pairs with unseen objects and compositions. Our code is available at https://diffusion-hyperfeatures.github.io.
title Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.14334