Affine-based Deformable Attention and Selective Fusion for Semi-dense Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hongkai, Luo, Zixin, Tian, Yurun, Bai, Xuyang, Wang, Ziyu, Zhou, Lei, Zhen, Mingmin, Fang, Tian, McKinnon, David, Tsin, Yanghai, Quan, Long
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911884756647936
author Chen, Hongkai
Luo, Zixin
Tian, Yurun
Bai, Xuyang
Wang, Ziyu
Zhou, Lei
Zhen, Mingmin
Fang, Tian
McKinnon, David
Tsin, Yanghai
Quan, Long
author_facet Chen, Hongkai
Luo, Zixin
Tian, Yurun
Bai, Xuyang
Wang, Ziyu
Zhou, Lei
Zhen, Mingmin
Fang, Tian
McKinnon, David
Tsin, Yanghai
Quan, Long
contents Identifying robust and accurate correspondences across images is a fundamental problem in computer vision that enables various downstream tasks. Recent semi-dense matching methods emphasize the effectiveness of fusing relevant cross-view information through Transformer. In this paper, we propose several improvements upon this paradigm. Firstly, we introduce affine-based local attention to model cross-view deformations. Secondly, we present selective fusion to merge local and global messages from cross attention. Apart from network structure, we also identify the importance of enforcing spatial smoothness in loss design, which has been omitted by previous works. Based on these augmentations, our network demonstrate strong matching capacity under different settings. The full version of our network achieves state-of-the-art performance among semi-dense matching methods at a similar cost to LoFTR, while the slim version reaches LoFTR baseline's performance with only 15% computation cost and 18% parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13874
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Affine-based Deformable Attention and Selective Fusion for Semi-dense Matching
Chen, Hongkai
Luo, Zixin
Tian, Yurun
Bai, Xuyang
Wang, Ziyu
Zhou, Lei
Zhen, Mingmin
Fang, Tian
McKinnon, David
Tsin, Yanghai
Quan, Long
Computer Vision and Pattern Recognition
Identifying robust and accurate correspondences across images is a fundamental problem in computer vision that enables various downstream tasks. Recent semi-dense matching methods emphasize the effectiveness of fusing relevant cross-view information through Transformer. In this paper, we propose several improvements upon this paradigm. Firstly, we introduce affine-based local attention to model cross-view deformations. Secondly, we present selective fusion to merge local and global messages from cross attention. Apart from network structure, we also identify the importance of enforcing spatial smoothness in loss design, which has been omitted by previous works. Based on these augmentations, our network demonstrate strong matching capacity under different settings. The full version of our network achieves state-of-the-art performance among semi-dense matching methods at a similar cost to LoFTR, while the slim version reaches LoFTR baseline's performance with only 15% computation cost and 18% parameters.
title Affine-based Deformable Attention and Selective Fusion for Semi-dense Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.13874