SegMASt3R: Geometry Grounded Segment Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jayanti, Rohit, Agrawal, Swayam, Garg, Vansh, Tourani, Siddharth, Khan, Muhammad Haris, Garg, Sourav, Krishna, Madhava
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908610214232064
author Jayanti, Rohit
Agrawal, Swayam
Garg, Vansh
Tourani, Siddharth
Khan, Muhammad Haris
Garg, Sourav
Krishna, Madhava
author_facet Jayanti, Rohit
Agrawal, Swayam
Garg, Vansh
Tourani, Siddharth
Khan, Muhammad Haris
Garg, Sourav
Krishna, Madhava
contents Segment matching is an important intermediate task in computer vision that establishes correspondences between semantically or geometrically coherent regions across images. Unlike keypoint matching, which focuses on localized features, segment matching captures structured regions, offering greater robustness to occlusions, lighting variations, and viewpoint changes. In this paper, we leverage the spatial understanding of 3D foundation models to tackle wide-baseline segment matching, a challenging setting involving extreme viewpoint shifts. We propose an architecture that uses the inductive bias of these 3D foundation models to match segments across image pairs with up to 180 degree view-point change rotation. Extensive experiments show that our approach outperforms state-of-the-art methods, including the SAM2 video propagator and local feature matching methods, by up to 30% on the AUPRC metric, on ScanNet++ and Replica datasets. We further demonstrate benefits of the proposed model on relevant downstream tasks, including 3D instance mapping and object-relative navigation. Project Page: https://segmast3r.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2510_05051
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SegMASt3R: Geometry Grounded Segment Matching
Jayanti, Rohit
Agrawal, Swayam
Garg, Vansh
Tourani, Siddharth
Khan, Muhammad Haris
Garg, Sourav
Krishna, Madhava
Computer Vision and Pattern Recognition
Segment matching is an important intermediate task in computer vision that establishes correspondences between semantically or geometrically coherent regions across images. Unlike keypoint matching, which focuses on localized features, segment matching captures structured regions, offering greater robustness to occlusions, lighting variations, and viewpoint changes. In this paper, we leverage the spatial understanding of 3D foundation models to tackle wide-baseline segment matching, a challenging setting involving extreme viewpoint shifts. We propose an architecture that uses the inductive bias of these 3D foundation models to match segments across image pairs with up to 180 degree view-point change rotation. Extensive experiments show that our approach outperforms state-of-the-art methods, including the SAM2 video propagator and local feature matching methods, by up to 30% on the AUPRC metric, on ScanNet++ and Replica datasets. We further demonstrate benefits of the proposed model on relevant downstream tasks, including 3D instance mapping and object-relative navigation. Project Page: https://segmast3r.github.io/
title SegMASt3R: Geometry Grounded Segment Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.05051