Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Junyi, Herrmann, Charles, Hur, Junhwa, Chen, Eric, Jampani, Varun, Sun, Deqing, Yang, Ming-Hsuan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911811296559104
author Zhang, Junyi
Herrmann, Charles
Hur, Junhwa
Chen, Eric
Jampani, Varun
Sun, Deqing
Yang, Ming-Hsuan
author_facet Zhang, Junyi
Herrmann, Charles
Hur, Junhwa
Chen, Eric
Jampani, Varun
Sun, Deqing
Yang, Ming-Hsuan
contents While pre-trained large-scale vision models have shown significant promise for semantic correspondence, their features often struggle to grasp the geometry and orientation of instances. This paper identifies the importance of being geometry-aware for semantic correspondence and reveals a limitation of the features of current foundation models under simple post-processing. We show that incorporating this information can markedly enhance semantic correspondence performance with simple but effective solutions in both zero-shot and supervised settings. We also construct a new challenging benchmark for semantic correspondence built from an existing animal pose estimation dataset, for both pre-training validating models. Our method achieves a PCK@0.10 score of 65.4 (zero-shot) and 85.6 (supervised) on the challenging SPair-71k dataset, outperforming the state of the art by 5.5p and 11.0p absolute gains, respectively. Our code and datasets are publicly available at: https://telling-left-from-right.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17034
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
Zhang, Junyi
Herrmann, Charles
Hur, Junhwa
Chen, Eric
Jampani, Varun
Sun, Deqing
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
While pre-trained large-scale vision models have shown significant promise for semantic correspondence, their features often struggle to grasp the geometry and orientation of instances. This paper identifies the importance of being geometry-aware for semantic correspondence and reveals a limitation of the features of current foundation models under simple post-processing. We show that incorporating this information can markedly enhance semantic correspondence performance with simple but effective solutions in both zero-shot and supervised settings. We also construct a new challenging benchmark for semantic correspondence built from an existing animal pose estimation dataset, for both pre-training validating models. Our method achieves a PCK@0.10 score of 65.4 (zero-shot) and 85.6 (supervised) on the challenging SPair-71k dataset, outperforming the state of the art by 5.5p and 11.0p absolute gains, respectively. Our code and datasets are publicly available at: https://telling-left-from-right.github.io/.
title Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.17034