CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Dongxu, Wang, Yingsen, Sun, Yiding, Xu, Haoran, Fan, Peilin, Zhu, Jihua
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914390596386816
author Zhang, Dongxu
Wang, Yingsen
Sun, Yiding
Xu, Haoran
Fan, Peilin
Zhu, Jihua
author_facet Zhang, Dongxu
Wang, Yingsen
Sun, Yiding
Xu, Haoran
Fan, Peilin
Zhu, Jihua
contents Robust point cloud registration is a fundamental task in 3D computer vision and geometric deep learning, essential for applications such as large-scale 3D reconstruction, augmented reality, and scene understanding. However, the performance of established learning-based methods often degrades in complex, real world scenarios characterized by incomplete data, sensor noise, and low overlap regions. To address these limitations, we propose CMHANet, a novel Cross-Modal Hybrid Attention Network. Our method integrates the fusion of rich contextual information from 2D images with the geometric detail of 3D point clouds, yielding a comprehensive and resilient feature representation. Furthermore, we introduce an innovative optimization function based on contrastive learning, which enforces geometric consistency and significantly improves the model's robustness to noise and partial observations. We evaluated CMHANet on the 3DMatch and the challenging 3DLoMatch datasets. \rev{Additionally, zero-shot evaluations on the TUM RGB-D SLAM dataset verify the model's generalization capability to unseen domains.} The experimental results demonstrate that our method achieves substantial improvements in both registration accuracy and overall robustness, outperforming current techniques. We also release our code in \href{https://github.com/DongXu-Zhang/CMHANet}{https://github.com/DongXu-Zhang/CMHANet}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12721
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration
Zhang, Dongxu
Wang, Yingsen
Sun, Yiding
Xu, Haoran
Fan, Peilin
Zhu, Jihua
Computer Vision and Pattern Recognition
Artificial Intelligence
Robust point cloud registration is a fundamental task in 3D computer vision and geometric deep learning, essential for applications such as large-scale 3D reconstruction, augmented reality, and scene understanding. However, the performance of established learning-based methods often degrades in complex, real world scenarios characterized by incomplete data, sensor noise, and low overlap regions. To address these limitations, we propose CMHANet, a novel Cross-Modal Hybrid Attention Network. Our method integrates the fusion of rich contextual information from 2D images with the geometric detail of 3D point clouds, yielding a comprehensive and resilient feature representation. Furthermore, we introduce an innovative optimization function based on contrastive learning, which enforces geometric consistency and significantly improves the model's robustness to noise and partial observations. We evaluated CMHANet on the 3DMatch and the challenging 3DLoMatch datasets. \rev{Additionally, zero-shot evaluations on the TUM RGB-D SLAM dataset verify the model's generalization capability to unseen domains.} The experimental results demonstrate that our method achieves substantial improvements in both registration accuracy and overall robustness, outperforming current techniques. We also release our code in \href{https://github.com/DongXu-Zhang/CMHANet}{https://github.com/DongXu-Zhang/CMHANet}.
title CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.12721