CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xuecong, Ding, Mengzhu, Sun, Zixuan, Li, Zhang, Teng, Xichao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908942477557760
author Liu, Xuecong
Ding, Mengzhu
Sun, Zixuan
Li, Zhang
Teng, Xichao
author_facet Liu, Xuecong
Ding, Mengzhu
Sun, Zixuan
Li, Zhang
Teng, Xichao
contents We present Consistent-Recurrent Feature Flow Transformer (CRFT), a unified coarse-to-fine framework based on feature flow learning for robust cross-modal image registration. CRFT learns a modality-independent feature flow representation within a transformer-based architecture that jointly performs feature alignment and flow estimation. The coarse stage establishes global correspondences through multi-scale feature correlation, while the fine stage refines local details via hierarchical feature fusion and adaptive spatial reasoning. To enhance geometric adaptability, an iterative discrepancy-guided attention mechanism with a Spatial Geometric Transform (SGT) recurrently refines the flow field, progressively capturing subtle spatial inconsistencies and enforcing feature-level consistency. This design enables accurate alignment under large affine and scale variations while maintaining structural coherence across modalities. Extensive experiments on diverse cross-modal datasets demonstrate that CRFT consistently outperforms state-of-the-art registration methods in both accuracy and robustness. Beyond registration, CRFT provides a generalizable paradigm for multimodal spatial correspondence, offering broad applicability to remote sensing, autonomous navigation, and medical imaging. Code and datasets are publicly available at https://github.com/NEU-Liuxuecong/CRFT.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05689
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration
Liu, Xuecong
Ding, Mengzhu
Sun, Zixuan
Li, Zhang
Teng, Xichao
Computer Vision and Pattern Recognition
Artificial Intelligence
We present Consistent-Recurrent Feature Flow Transformer (CRFT), a unified coarse-to-fine framework based on feature flow learning for robust cross-modal image registration. CRFT learns a modality-independent feature flow representation within a transformer-based architecture that jointly performs feature alignment and flow estimation. The coarse stage establishes global correspondences through multi-scale feature correlation, while the fine stage refines local details via hierarchical feature fusion and adaptive spatial reasoning. To enhance geometric adaptability, an iterative discrepancy-guided attention mechanism with a Spatial Geometric Transform (SGT) recurrently refines the flow field, progressively capturing subtle spatial inconsistencies and enforcing feature-level consistency. This design enables accurate alignment under large affine and scale variations while maintaining structural coherence across modalities. Extensive experiments on diverse cross-modal datasets demonstrate that CRFT consistently outperforms state-of-the-art registration methods in both accuracy and robustness. Beyond registration, CRFT provides a generalizable paradigm for multimodal spatial correspondence, offering broad applicability to remote sensing, autonomous navigation, and medical imaging. Code and datasets are publicly available at https://github.com/NEU-Liuxuecong/CRFT.
title CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.05689