MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xingyi, Yu, Hao, Peng, Sida, Tan, Dongli, Shen, Zehong, Bao, Hujun, Zhou, Xiaowei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Generating Human Motion in 3D Scenes from Text Descriptions
by: Cen, Zhi, et al.
Published: (2024)
by: Cen, Zhi, et al.
Published: (2024)
World-Grounded Human Motion Recovery via Gravity-View Coordinates
by: Shen, Zehong, et al.
Published: (2024)
by: Shen, Zehong, et al.
Published: (2024)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024)
by: Lin, Haotong, et al.
Published: (2024)
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
by: Xu, Zhen, et al.
Published: (2024)
by: Xu, Zhen, et al.
Published: (2024)
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Dyn-E: Local Appearance Editing of Dynamic Neural Radiance Fields
by: Zhang, Shangzan, et al.
Published: (2023)
by: Zhang, Shangzan, et al.
Published: (2023)
Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation
by: Cen, Zhi, et al.
Published: (2025)
by: Cen, Zhi, et al.
Published: (2025)
UFM: Unified Feature Matching Pre-training with Multi-Modal Image Assistants
by: Di, Yide, et al.
Published: (2025)
by: Di, Yide, et al.
Published: (2025)
MATCHA:Towards Matching Anything
by: Xue, Fei, et al.
Published: (2025)
by: Xue, Fei, et al.
Published: (2025)
Stereo Anything: Unifying Zero-shot Stereo Matching with Large-Scale Mixed Data
by: Guo, Xianda, et al.
Published: (2024)
by: Guo, Xianda, et al.
Published: (2024)
UDuo: Universal Dual Optimization Framework for Online Matching
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
EnvGS: Modeling View-Dependent Appearance with Environment Gaussian
by: Xie, Tao, et al.
Published: (2024)
by: Xie, Tao, et al.
Published: (2024)
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
by: Jin, Yudong, et al.
Published: (2025)
by: Jin, Yudong, et al.
Published: (2025)
Precise Action-to-Video Generation Through Visual Action Prompts
by: Wang, Yuang, et al.
Published: (2025)
by: Wang, Yuang, et al.
Published: (2025)
Multi-view Reconstruction via SfM-guided Monocular Depth Estimation
by: Guo, Haoyu, et al.
Published: (2025)
by: Guo, Haoyu, et al.
Published: (2025)
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
MINIMA: Modality Invariant Image Matching
by: Ren, Jiangwei, et al.
Published: (2024)
by: Ren, Jiangwei, et al.
Published: (2024)
IntrinsicAnything: Learning Diffusion Priors for Inverse Rendering Under Unknown Illumination
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Contrast-X: A Multi-Modal Contrast Image Synthesis Benchmark and Universal Modality Flow Matching
by: Chen, Yifan, et al.
Published: (2026)
by: Chen, Yifan, et al.
Published: (2026)
MaPa: Text-driven Photorealistic Material Painting for 3D Shapes
by: Zhang, Shangzan, et al.
Published: (2024)
by: Zhang, Shangzan, et al.
Published: (2024)
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation
by: Yu, Yuanhong, et al.
Published: (2025)
by: Yu, Yuanhong, et al.
Published: (2025)
Adaptive Multi-Modal Cross-Entropy Loss for Stereo Matching
by: Xu, Peng, et al.
Published: (2023)
by: Xu, Peng, et al.
Published: (2023)
StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
FreeTimeGS: Free Gaussian Primitives at Anytime and Anywhere for Dynamic Scene Reconstruction
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
MESA: Matching Everything by Segmenting Anything
by: Zhang, Yesheng, et al.
Published: (2024)
by: Zhang, Yesheng, et al.
Published: (2024)
Automatic Creative Selection with Cross-Modal Matching
by: Kim, Alex, et al.
Published: (2024)
by: Kim, Alex, et al.
Published: (2024)
HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
by: Yu, Zhiyuan, et al.
Published: (2025)
by: Yu, Zhiyuan, et al.
Published: (2025)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
by: Han, Yu, et al.
Published: (2025)
by: Han, Yu, et al.
Published: (2025)
PATS: Patch Area Transportation with Subdivision for Local Feature Matching
by: Ni, Junjie, et al.
Published: (2023)
by: Ni, Junjie, et al.
Published: (2023)
Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
by: Liu, Weide, et al.
Published: (2025)
by: Liu, Weide, et al.
Published: (2025)
Class-Agnostic Region-of-Interest Matching in Document Images
by: Zhang, Demin, et al.
Published: (2025)
by: Zhang, Demin, et al.
Published: (2025)
Cross-Modal Entity Matching for Visually Rich Documents
by: Sarkhel, Ritesh, et al.
Published: (2023)
by: Sarkhel, Ritesh, et al.
Published: (2023)
CM-Bench: A Comprehensive Cross-Modal Feature Matching Benchmark Bridging Visible and Infrared Images
by: Sun, Liangzheng, et al.
Published: (2026)
by: Sun, Liangzheng, et al.
Published: (2026)
Boosting Image Restoration via Priors from Pre-trained Models
by: Xu, Xiaogang, et al.
Published: (2024)
by: Xu, Xiaogang, et al.
Published: (2024)
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
by: Xie, Tao, et al.
Published: (2026)
by: Xie, Tao, et al.
Published: (2026)
Generalized Large-Scale Data Condensation via Various Backbone and Statistical Matching
by: Shao, Shitong, et al.
Published: (2023)
by: Shao, Shitong, et al.
Published: (2023)
Similar Items
-
Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like Speed
by: Wang, Yifan, et al.
Published: (2024) -
Generating Human Motion in 3D Scenes from Text Descriptions
by: Cen, Zhi, et al.
Published: (2024) -
World-Grounded Human Motion Recovery via Gravity-View Coordinates
by: Shen, Zehong, et al.
Published: (2024) -
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024) -
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025)