UFM: Unified Feature Matching Pre-training with Multi-Modal Image Assistants

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Di, Yide, Liao, Yun, Zhou, Hao, Zhu, Kaijun, Duan, Qing, Liu, Junhui, Lu, Mingyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908288156696576
author Di, Yide
Liao, Yun
Zhou, Hao
Zhu, Kaijun
Duan, Qing
Liu, Junhui
Lu, Mingyu
author_facet Di, Yide
Liao, Yun
Zhou, Hao
Zhu, Kaijun
Duan, Qing
Liu, Junhui
Lu, Mingyu
contents Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Feature Matching pre-trained model (UFM) designed to address feature matching challenges across a wide spectrum of modal images. We present Multimodal Image Assistant (MIA) transformers, finely tunable structures adept at handling diverse feature matching problems. UFM exhibits versatility in addressing both feature matching tasks within the same modal and those across different modals. Additionally, we propose a data augmentation algorithm and a staged pre-training strategy to effectively tackle challenges arising from sparse data in specific modals and imbalanced modal datasets. Experimental results demonstrate that UFM excels in generalization and performance across various feature matching tasks. The code will be released at:https://github.com/LiaoYun0x0/UFM.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UFM: Unified Feature Matching Pre-training with Multi-Modal Image Assistants
Di, Yide
Liao, Yun
Zhou, Hao
Zhu, Kaijun
Duan, Qing
Liu, Junhui
Lu, Mingyu
Computer Vision and Pattern Recognition
Image and Video Processing
Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Feature Matching pre-trained model (UFM) designed to address feature matching challenges across a wide spectrum of modal images. We present Multimodal Image Assistant (MIA) transformers, finely tunable structures adept at handling diverse feature matching problems. UFM exhibits versatility in addressing both feature matching tasks within the same modal and those across different modals. Additionally, we propose a data augmentation algorithm and a staged pre-training strategy to effectively tackle challenges arising from sparse data in specific modals and imbalanced modal datasets. Experimental results demonstrate that UFM excels in generalization and performance across various feature matching tasks. The code will be released at:https://github.com/LiaoYun0x0/UFM.
title UFM: Unified Feature Matching Pre-training with Multi-Modal Image Assistants
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2503.21820