MARCO: Navigating the Unseen Space of Semantic Correspondence

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cuttano, Claudia, Trivigno, Gabriele, Masone, Carlo, Roth, Stefan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913047667277824
author Cuttano, Claudia
Trivigno, Gabriele
Masone, Carlo
Roth, Stefan
author_facet Cuttano, Claudia
Trivigno, Gabriele
Masone, Carlo
Roth, Stefan
contents Recent advances in semantic correspondence rely on dual-encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion-parameter models generalize poorly beyond training keypoints, revealing a gap between benchmark performance and real-world usability, where queried points rarely match those seen during training. Building upon DINOv2, we introduce MARCO, a unified model for generalizable correspondence driven by a novel training framework that enhances both fine-grained localization and semantic generalization. By coupling a coarse-to-fine objective that refines spatial precision with a self-distillation framework, which expands sparse supervision beyond annotated regions, our approach transforms a handful of keypoints into dense, semantically coherent correspondences. MARCO sets a new state of the art on SPair-71k, AP-10K, and PF-PASCAL, with gains that amplify at fine-grained localization thresholds (+8.9 PCK@0.01), strongest generalization to unseen keypoints (+5.1, SPair-U) and categories (+4.7, MP-100), while remaining 3x smaller and 10x faster than diffusion-based approaches. Code is available at https://github.com/visinf/MARCO .
format Preprint
id arxiv_https___arxiv_org_abs_2604_18267
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MARCO: Navigating the Unseen Space of Semantic Correspondence
Cuttano, Claudia
Trivigno, Gabriele
Masone, Carlo
Roth, Stefan
Computer Vision and Pattern Recognition
Recent advances in semantic correspondence rely on dual-encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion-parameter models generalize poorly beyond training keypoints, revealing a gap between benchmark performance and real-world usability, where queried points rarely match those seen during training. Building upon DINOv2, we introduce MARCO, a unified model for generalizable correspondence driven by a novel training framework that enhances both fine-grained localization and semantic generalization. By coupling a coarse-to-fine objective that refines spatial precision with a self-distillation framework, which expands sparse supervision beyond annotated regions, our approach transforms a handful of keypoints into dense, semantically coherent correspondences. MARCO sets a new state of the art on SPair-71k, AP-10K, and PF-PASCAL, with gains that amplify at fine-grained localization thresholds (+8.9 PCK@0.01), strongest generalization to unseen keypoints (+5.1, SPair-U) and categories (+4.7, MP-100), while remaining 3x smaller and 10x faster than diffusion-based approaches. Code is available at https://github.com/visinf/MARCO .
title MARCO: Navigating the Unseen Space of Semantic Correspondence
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.18267