Self-Supervised Spatial Correspondence Across Modalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shrivastava, Ayush, Owens, Andrew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909635261235200
author Shrivastava, Ayush
Owens, Andrew
author_facet Shrivastava, Ayush
Owens, Andrew
contents We present a method for finding cross-modal space-time correspondences. Given two images from different visual modalities, such as an RGB image and a depth map, our model identifies which pairs of pixels correspond to the same physical points in the scene. To solve this problem, we extend the contrastive random walk framework to simultaneously learn cycle-consistent feature representations for both cross-modal and intra-modal matching. The resulting model is simple and has no explicit photo-consistency assumptions. It can be trained entirely using unlabeled data, without the need for any spatially aligned multimodal image pairs. We evaluate our method on both geometric and semantic correspondence tasks. For geometric matching, we consider challenging tasks such as RGB-to-depth and RGB-to-thermal matching (and vice versa); for semantic matching, we evaluate on photo-sketch and cross-style image alignment. Our method achieves strong performance across all benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03148
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Supervised Spatial Correspondence Across Modalities
Shrivastava, Ayush
Owens, Andrew
Computer Vision and Pattern Recognition
We present a method for finding cross-modal space-time correspondences. Given two images from different visual modalities, such as an RGB image and a depth map, our model identifies which pairs of pixels correspond to the same physical points in the scene. To solve this problem, we extend the contrastive random walk framework to simultaneously learn cycle-consistent feature representations for both cross-modal and intra-modal matching. The resulting model is simple and has no explicit photo-consistency assumptions. It can be trained entirely using unlabeled data, without the need for any spatially aligned multimodal image pairs. We evaluate our method on both geometric and semantic correspondence tasks. For geometric matching, we consider challenging tasks such as RGB-to-depth and RGB-to-thermal matching (and vice versa); for semantic matching, we evaluate on photo-sketch and cross-style image alignment. Our method achieves strong performance across all benchmarks.
title Self-Supervised Spatial Correspondence Across Modalities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.03148