Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: You, Yang, Li, Yixin, Deng, Congyue, Wang, Yue, Guibas, Leonidas
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910835319767040
author You, Yang
Li, Yixin
Deng, Congyue
Wang, Yue
Guibas, Leonidas
author_facet You, Yang
Li, Yixin
Deng, Congyue
Wang, Yue
Guibas, Leonidas
contents Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D awareness of ViT-based models. We begin by systematically assessing their ability to learn 3D equivariant features, specifically examining the consistency of semantic embeddings across different viewpoints. Our findings indicate that improved 3D equivariance leads to better performance on various downstream tasks, including pose estimation, tracking, and semantic transfer. Building on this insight, we propose a simple yet effective finetuning strategy based on 3D correspondences, which significantly enhances the 3D correspondence understanding of existing vision models. Remarkably, finetuning on a single object for one iteration results in substantial gains. Our code is available at https://github.com/qq456cvb/3DCorrEnhance.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19458
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
You, Yang
Li, Yixin
Deng, Congyue
Wang, Yue
Guibas, Leonidas
Computer Vision and Pattern Recognition
Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D awareness of ViT-based models. We begin by systematically assessing their ability to learn 3D equivariant features, specifically examining the consistency of semantic embeddings across different viewpoints. Our findings indicate that improved 3D equivariance leads to better performance on various downstream tasks, including pose estimation, tracking, and semantic transfer. Building on this insight, we propose a simple yet effective finetuning strategy based on 3D correspondences, which significantly enhances the 3D correspondence understanding of existing vision models. Remarkably, finetuning on a single object for one iteration results in substantial gains. Our code is available at https://github.com/qq456cvb/3DCorrEnhance.
title Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.19458