ETO:Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography Hypotheses
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915098278232064 |
|---|---|
| author | Ni, Junjie Zhang, Guofeng Li, Guanglin Li, Yijin Liu, Xinyang Huang, Zhaoyang Bao, Hujun |
| author_facet | Ni, Junjie Zhang, Guofeng Li, Guanglin Li, Yijin Liu, Xinyang Huang, Zhaoyang Bao, Hujun |
| contents | We tackle the efficiency problem of learning local feature matching. Recent advancements have given rise to purely CNN-based and transformer-based approaches, each augmented with deep learning techniques. While CNN-based methods often excel in matching speed, transformer-based methods tend to provide more accurate matches. We propose an efficient transformer-based network architecture for local feature matching. This technique is built on constructing multiple homography hypotheses to approximate the continuous correspondence in the real world and uni-directional cross-attention to accelerate the refinement. On the YFCC100M dataset, our matching accuracy is competitive with LoFTR, a state-of-the-art transformer-based architecture, while the inference speed is boosted to 4 times, even outperforming the CNN-based methods. Comprehensive evaluations on other open datasets such as Megadepth, ScanNet, and HPatches demonstrate our method's efficacy, highlighting its potential to significantly enhance a wide array of downstream applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_22733 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ETO:Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography Hypotheses Ni, Junjie Zhang, Guofeng Li, Guanglin Li, Yijin Liu, Xinyang Huang, Zhaoyang Bao, Hujun Computer Vision and Pattern Recognition We tackle the efficiency problem of learning local feature matching. Recent advancements have given rise to purely CNN-based and transformer-based approaches, each augmented with deep learning techniques. While CNN-based methods often excel in matching speed, transformer-based methods tend to provide more accurate matches. We propose an efficient transformer-based network architecture for local feature matching. This technique is built on constructing multiple homography hypotheses to approximate the continuous correspondence in the real world and uni-directional cross-attention to accelerate the refinement. On the YFCC100M dataset, our matching accuracy is competitive with LoFTR, a state-of-the-art transformer-based architecture, while the inference speed is boosted to 4 times, even outperforming the CNN-based methods. Comprehensive evaluations on other open datasets such as Megadepth, ScanNet, and HPatches demonstrate our method's efficacy, highlighting its potential to significantly enhance a wide array of downstream applications. |
| title | ETO:Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography Hypotheses |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.22733 |