RoMa: Robust Dense Feature Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Edstedt, Johan, Sun, Qiyu, Bökman, Georg, Wadenbäck, Mårten, Felsberg, Michael
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908955229290496
author Edstedt, Johan
Sun, Qiyu
Bökman, Georg
Wadenbäck, Mårten
Felsberg, Michael
author_facet Edstedt, Johan
Sun, Qiyu
Bökman, Georg
Wadenbäck, Mårten
Felsberg, Michael
contents Feature matching is an important computer vision task that involves estimating correspondences between two images of a 3D scene, and dense methods estimate all such correspondences. The aim is to learn a robust model, i.e., a model able to match under challenging real-world changes. In this work, we propose such a model, leveraging frozen pretrained features from the foundation model DINOv2. Although these features are significantly more robust than local features trained from scratch, they are inherently coarse. We therefore combine them with specialized ConvNet fine features, creating a precisely localizable feature pyramid. To further improve robustness, we propose a tailored transformer match decoder that predicts anchor probabilities, which enables it to express multimodality. Finally, we propose an improved loss formulation through regression-by-classification with subsequent robust regression. We conduct a comprehensive set of experiments that show that our method, RoMa, achieves significant gains, setting a new state-of-the-art. In particular, we achieve a 36% improvement on the extremely challenging WxBS benchmark. Code is provided at https://github.com/Parskatt/RoMa
format Preprint
id arxiv_https___arxiv_org_abs_2305_15404
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle RoMa: Robust Dense Feature Matching
Edstedt, Johan
Sun, Qiyu
Bökman, Georg
Wadenbäck, Mårten
Felsberg, Michael
Computer Vision and Pattern Recognition
Feature matching is an important computer vision task that involves estimating correspondences between two images of a 3D scene, and dense methods estimate all such correspondences. The aim is to learn a robust model, i.e., a model able to match under challenging real-world changes. In this work, we propose such a model, leveraging frozen pretrained features from the foundation model DINOv2. Although these features are significantly more robust than local features trained from scratch, they are inherently coarse. We therefore combine them with specialized ConvNet fine features, creating a precisely localizable feature pyramid. To further improve robustness, we propose a tailored transformer match decoder that predicts anchor probabilities, which enables it to express multimodality. Finally, we propose an improved loss formulation through regression-by-classification with subsequent robust regression. We conduct a comprehensive set of experiments that show that our method, RoMa, achieves significant gains, setting a new state-of-the-art. In particular, we achieve a 36% improvement on the extremely challenging WxBS benchmark. Code is provided at https://github.com/Parskatt/RoMa
title RoMa: Robust Dense Feature Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.15404