Mind the Gap: Aligning Vision Foundation Models to Image Feature Matching
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yuhan, Fu, Jingwen, Wu, Yang, Wu, Kangyi, Li, Pengna, Wu, Jiayi, Zhou, Sanping, Xin, Jingmin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
REGNav: Room Expert Guided Image-Goal Navigation
di: Li, Pengna, et al.
Pubblicazione: (2025)
di: Li, Pengna, et al.
Pubblicazione: (2025)
Camera-aware Label Refinement for Unsupervised Person Re-identification
di: Li, Pengna, et al.
Pubblicazione: (2024)
di: Li, Pengna, et al.
Pubblicazione: (2024)
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
di: Wu, Kangyi, et al.
Pubblicazione: (2026)
di: Wu, Kangyi, et al.
Pubblicazione: (2026)
SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation
di: Li, Pengna, et al.
Pubblicazione: (2026)
di: Li, Pengna, et al.
Pubblicazione: (2026)
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
di: Wu, Kangyi, et al.
Pubblicazione: (2025)
di: Wu, Kangyi, et al.
Pubblicazione: (2025)
Aligning Neuronal Coding of Dynamic Visual Scenes with Foundation Vision Models
di: Wu, Rining, et al.
Pubblicazione: (2024)
di: Wu, Rining, et al.
Pubblicazione: (2024)
Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
PromptMID: Modal Invariant Descriptors Based on Diffusion and Vision Foundation Models for Optical-SAR Image Matching
di: Nie, Han, et al.
Pubblicazione: (2025)
di: Nie, Han, et al.
Pubblicazione: (2025)
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
di: Lyu, Kailin, et al.
Pubblicazione: (2026)
di: Lyu, Kailin, et al.
Pubblicazione: (2026)
DAMap: Distance-aware MapNet for High Quality HD Map Construction
di: Dong, Jinpeng, et al.
Pubblicazione: (2025)
di: Dong, Jinpeng, et al.
Pubblicazione: (2025)
StructVPR++: Distill Structural and Semantic Knowledge with Weighting Samples for Visual Place Recognition
di: Shen, Yanqing, et al.
Pubblicazione: (2025)
di: Shen, Yanqing, et al.
Pubblicazione: (2025)
DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching
di: Yang, Meng, et al.
Pubblicazione: (2025)
di: Yang, Meng, et al.
Pubblicazione: (2025)
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
di: Wu, Yuxuan, et al.
Pubblicazione: (2025)
di: Wu, Yuxuan, et al.
Pubblicazione: (2025)
The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation
di: Liu, Zhen, et al.
Pubblicazione: (2026)
di: Liu, Zhen, et al.
Pubblicazione: (2026)
Bridge the Gap between SNN and ANN for Image Restoration
di: Su, Xin, et al.
Pubblicazione: (2025)
di: Su, Xin, et al.
Pubblicazione: (2025)
FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation
di: Wang, Sen, et al.
Pubblicazione: (2025)
di: Wang, Sen, et al.
Pubblicazione: (2025)
SAMPO-Path: Segmentation Intent-Aligned Preference Optimization for Pathology Foundation Model Segmentation
di: Wu, Yonghuang, et al.
Pubblicazione: (2025)
di: Wu, Yonghuang, et al.
Pubblicazione: (2025)
Multi-modal Attribute Prompting for Vision-Language Models
di: Liu, Xin, et al.
Pubblicazione: (2024)
di: Liu, Xin, et al.
Pubblicazione: (2024)
Semantic-aware Representation Learning for Homography Estimation
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
CosmicMan: A Text-to-Image Foundation Model for Humans
di: Li, Shikai, et al.
Pubblicazione: (2024)
di: Li, Shikai, et al.
Pubblicazione: (2024)
VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
di: Sargent, Kyle, et al.
Pubblicazione: (2025)
di: Sargent, Kyle, et al.
Pubblicazione: (2025)
All-in-One: Transferring Vision Foundation Models into Stereo Matching
di: Zhou, Jingyi, et al.
Pubblicazione: (2024)
di: Zhou, Jingyi, et al.
Pubblicazione: (2024)
Recurrent Aligned Network for Generalized Pedestrian Trajectory Prediction
di: Dong, Yonghao, et al.
Pubblicazione: (2024)
di: Dong, Yonghao, et al.
Pubblicazione: (2024)
Fusion of Foundation and Vision Transformer Model Features for Dermatoscopic Image Classification
di: Mahbod, Amirreza, et al.
Pubblicazione: (2025)
di: Mahbod, Amirreza, et al.
Pubblicazione: (2025)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
di: Jiang, Songtao, et al.
Pubblicazione: (2025)
di: Jiang, Songtao, et al.
Pubblicazione: (2025)
Mind the Gap: Continuous Magnification Sampling for Pathology Foundation Models
di: Möllers, Alexander, et al.
Pubblicazione: (2026)
di: Möllers, Alexander, et al.
Pubblicazione: (2026)
Categorical Knowledge Fused Recognition: Fusing Hierarchical Knowledge with Image Classification through Aligning and Deep Metric Learning
di: Zhao, Yunfeng, et al.
Pubblicazione: (2024)
di: Zhao, Yunfeng, et al.
Pubblicazione: (2024)
LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
di: Liang, Zhanhao, et al.
Pubblicazione: (2026)
di: Liang, Zhanhao, et al.
Pubblicazione: (2026)
Advancing Video Self-Supervised Learning via Image Foundation Models
di: Wu, Jingwei, et al.
Pubblicazione: (2025)
di: Wu, Jingwei, et al.
Pubblicazione: (2025)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
di: Chen, Zhe, et al.
Pubblicazione: (2023)
di: Chen, Zhe, et al.
Pubblicazione: (2023)
MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise
di: Wu, Ruiqi, et al.
Pubblicazione: (2024)
di: Wu, Ruiqi, et al.
Pubblicazione: (2024)
Leveraging Diffusion Model and Image Foundation Model for Improved Correspondence Matching in Coronary Angiography
di: Zhao, Lin, et al.
Pubblicazione: (2025)
di: Zhao, Lin, et al.
Pubblicazione: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
Mind the Gap Between Prototypes and Images in Cross-domain Finetuning
di: Tian, Hongduan, et al.
Pubblicazione: (2024)
di: Tian, Hongduan, et al.
Pubblicazione: (2024)
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
di: Xiao, Yao, et al.
Pubblicazione: (2025)
di: Xiao, Yao, et al.
Pubblicazione: (2025)
Advancing Pre-trained Teacher: Towards Robust Feature Discrepancy for Anomaly Detection
di: Tang, Canhui, et al.
Pubblicazione: (2024)
di: Tang, Canhui, et al.
Pubblicazione: (2024)
PMT: Progressive Mean Teacher via Exploring Temporal Consistency for Semi-Supervised Medical Image Segmentation
di: Gao, Ning, et al.
Pubblicazione: (2024)
di: Gao, Ning, et al.
Pubblicazione: (2024)
Labeled-to-Unlabeled Distribution Alignment for Partially-Supervised Multi-Organ Medical Image Segmentation
di: Jiang, Xixi, et al.
Pubblicazione: (2024)
di: Jiang, Xixi, et al.
Pubblicazione: (2024)
VisionCLIP: An Med-AIGC based Ethical Language-Image Foundation Model for Generalizable Retina Image Analysis
di: Wei, Hao, et al.
Pubblicazione: (2024)
di: Wei, Hao, et al.
Pubblicazione: (2024)
Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
Documenti analoghi
-
REGNav: Room Expert Guided Image-Goal Navigation
di: Li, Pengna, et al.
Pubblicazione: (2025) -
Camera-aware Label Refinement for Unsupervised Person Re-identification
di: Li, Pengna, et al.
Pubblicazione: (2024) -
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
di: Wu, Kangyi, et al.
Pubblicazione: (2026) -
SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation
di: Li, Pengna, et al.
Pubblicazione: (2026) -
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
di: Wu, Kangyi, et al.
Pubblicazione: (2025)