OSDM-MReg: Multimodal Image Registration based One Step Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Xiaochen, Guo, Weiwei, Yu, Wenxian, Wei, Feiming, Li, Dongying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918361996197888
author Wei, Xiaochen
Guo, Weiwei
Yu, Wenxian
Wei, Feiming
Li, Dongying
author_facet Wei, Xiaochen
Guo, Weiwei
Yu, Wenxian
Wei, Feiming
Li, Dongying
contents Multimodal remote sensing image registration aligns images from different sensors for data fusion and analysis. However, existing methods often struggle to extract modality-invariant features when faced with large nonlinear radiometric differences, such as those between SAR and optical images. To address these challenges, we propose OSDM-MReg, a novel multimodal image registration framework that bridges the modality gap through image-to-image translation. Specifically, we introduce a one-step unaligned target-guided conditional diffusion model (UTGOS-CDM) to translate source and target images into a unified representation domain. Unlike traditional conditional DDPM that require hundreds of iterative steps for inference, our model incorporates a novel inverse translation objective during training to enable direct prediction of the translated image in a single step at test time, significantly accelerating the registration process. After translation, we design a multimodal multiscale registration network (MM-Reg) that extracts and fuses both unimodal and translated multimodal images using the proposed multimodal fusion strategy, enhancing the robustness and precision of alignment across scales and modalities. Extensive experiments on the OSdataset demonstrate that OSDM-MReg achieves superior registration accuracy compared to state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OSDM-MReg: Multimodal Image Registration based One Step Diffusion Model
Wei, Xiaochen
Guo, Weiwei
Yu, Wenxian
Wei, Feiming
Li, Dongying
Computer Vision and Pattern Recognition
Image and Video Processing
Multimodal remote sensing image registration aligns images from different sensors for data fusion and analysis. However, existing methods often struggle to extract modality-invariant features when faced with large nonlinear radiometric differences, such as those between SAR and optical images. To address these challenges, we propose OSDM-MReg, a novel multimodal image registration framework that bridges the modality gap through image-to-image translation. Specifically, we introduce a one-step unaligned target-guided conditional diffusion model (UTGOS-CDM) to translate source and target images into a unified representation domain. Unlike traditional conditional DDPM that require hundreds of iterative steps for inference, our model incorporates a novel inverse translation objective during training to enable direct prediction of the translated image in a single step at test time, significantly accelerating the registration process. After translation, we design a multimodal multiscale registration network (MM-Reg) that extracts and fuses both unimodal and translated multimodal images using the proposed multimodal fusion strategy, enhancing the robustness and precision of alignment across scales and modalities. Extensive experiments on the OSdataset demonstrate that OSDM-MReg achieves superior registration accuracy compared to state-of-the-art methods.
title OSDM-MReg: Multimodal Image Registration based One Step Diffusion Model
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2504.06027