Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Muyang, Liu, Yucheng, Ma, Jianbo, Osborne, Elliot, Han, Bo, Liu, Tongliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gromov-Wasserstein-like Distances in the Gaussian Mixture Models Space
by: Salmona, Antoine, et al.
Published: (2023)
by: Salmona, Antoine, et al.
Published: (2023)
Min Generalized Sliced Gromov Wasserstein: A Scalable Path to Gromov Wasserstein
by: Shahbazi, Ashkan, et al.
Published: (2026)
by: Shahbazi, Ashkan, et al.
Published: (2026)
Shape-of-You: Fused Gromov-Wasserstein Optimal Transport for Semantic Correspondence in-the-Wild
by: Im, Jiin, et al.
Published: (2026)
by: Im, Jiin, et al.
Published: (2026)
Improving Hyperbolic Representations via Gromov-Wasserstein Regularization
by: Yang, Yifei, et al.
Published: (2024)
by: Yang, Yifei, et al.
Published: (2024)
MOKD: Cross-domain Finetuning for Few-shot Classification via Maximizing Optimized Kernel Dependence
by: Tian, Hongduan, et al.
Published: (2024)
by: Tian, Hongduan, et al.
Published: (2024)
Negative Label Guided OOD Detection with Pretrained Vision-Language Models
by: Jiang, Xue, et al.
Published: (2024)
by: Jiang, Xue, et al.
Published: (2024)
Mind the Gap Between Prototypes and Images in Cross-domain Finetuning
by: Tian, Hongduan, et al.
Published: (2024)
by: Tian, Hongduan, et al.
Published: (2024)
BadLabel: A Robust Perspective on Evaluating and Enhancing Label-noise Learning
by: Zhang, Jingfeng, et al.
Published: (2023)
by: Zhang, Jingfeng, et al.
Published: (2023)
Stereographic Spherical Sliced Wasserstein Distances
by: Tran, Huy, et al.
Published: (2024)
by: Tran, Huy, et al.
Published: (2024)
Energy-Based Sliced Wasserstein Distance
by: Nguyen, Khai, et al.
Published: (2023)
by: Nguyen, Khai, et al.
Published: (2023)
Max-Sliced Wasserstein Distance and its use for GANs
by: Deshpande, Ishan, et al.
Published: (2019)
by: Deshpande, Ishan, et al.
Published: (2019)
Rethinking Meta-Learning from a Learning Lens
by: Wang, Jingyao, et al.
Published: (2024)
by: Wang, Jingyao, et al.
Published: (2024)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
by: Wu, Zhenkai, et al.
Published: (2025)
by: Wu, Zhenkai, et al.
Published: (2025)
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026)
by: Lin, Weifeng, et al.
Published: (2026)
Large-scale Dataset Pruning with Dynamic Uncertainty
by: He, Muyang, et al.
Published: (2023)
by: He, Muyang, et al.
Published: (2023)
Few-Shot Adversarial Prompt Learning on Vision-Language Models
by: Zhou, Yiwei, et al.
Published: (2024)
by: Zhou, Yiwei, et al.
Published: (2024)
Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge Distillation
by: Lv, Jiaming, et al.
Published: (2024)
by: Lv, Jiaming, et al.
Published: (2024)
Privacy-Preserving Model Transcription with Differentially Private Synthetic Distillation
by: Liu, Bochao, et al.
Published: (2026)
by: Liu, Bochao, et al.
Published: (2026)
Beyond Perceptual Distances: Rethinking Disparity Assessment for Out-of-Distribution Detection with Diffusion Models
by: Fang, Kun, et al.
Published: (2024)
by: Fang, Kun, et al.
Published: (2024)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
by: Liu, Jiajin, et al.
Published: (2026)
by: Liu, Jiajin, et al.
Published: (2026)
Improving Accuracy-robustness Trade-off via Pixel Reweighted Adversarial Training
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
by: Bai, Andrew, et al.
Published: (2025)
by: Bai, Andrew, et al.
Published: (2025)
SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
by: Yi, Huahui, et al.
Published: (2025)
by: Yi, Huahui, et al.
Published: (2025)
Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
by: Shi, Yucheng, et al.
Published: (2025)
by: Shi, Yucheng, et al.
Published: (2025)
ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
by: Li, Junxian, et al.
Published: (2024)
by: Li, Junxian, et al.
Published: (2024)
DocVLM: Make Your VLM an Efficient Reader
by: Nacson, Mor Shpigel, et al.
Published: (2024)
by: Nacson, Mor Shpigel, et al.
Published: (2024)
When Data-Free Knowledge Distillation Meets Non-Transferable Teacher: Escaping Out-of-Distribution Trap is All You Need
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
Toward Robust Non-Transferable Learning: A Survey and Benchmark
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
Rethinking Weight-Averaged Model-merging
by: Wang, Hu, et al.
Published: (2024)
by: Wang, Hu, et al.
Published: (2024)
Rethinking Encoder-Decoder Flow Through Shared Structures
by: Laboyrie, Frederik, et al.
Published: (2025)
by: Laboyrie, Frederik, et al.
Published: (2025)
Disentangled Representation Learning with the Gromov-Monge Gap
by: Uscidda, Théo, et al.
Published: (2024)
by: Uscidda, Théo, et al.
Published: (2024)
Wasserstein Distances Made Explainable: Insights Into Dataset Shifts and Transport Phenomena
by: Naumann, Philip, et al.
Published: (2025)
by: Naumann, Philip, et al.
Published: (2025)
Omnimodal Dataset Distillation via High-order Proxy Alignment
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
Architecture, Dataset and Model-Scale Agnostic Data-free Meta-Learning
by: Hu, Zixuan, et al.
Published: (2023)
by: Hu, Zixuan, et al.
Published: (2023)
SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
by: Li, Muyang, et al.
Published: (2024)
by: Li, Muyang, et al.
Published: (2024)
Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
by: Zhong, Siru, et al.
Published: (2025)
by: Zhong, Siru, et al.
Published: (2025)
SubFlow: Sub-mode Conditioned Flow Matching for Diverse One-Step Generation
by: Lin, Yexiong, et al.
Published: (2026)
by: Lin, Yexiong, et al.
Published: (2026)
Rethinking Genomic Modeling Through Optical Character Recognition
by: Xiang, Hongxin, et al.
Published: (2026)
by: Xiang, Hongxin, et al.
Published: (2026)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
by: Bao, Chen, et al.
Published: (2024)
by: Bao, Chen, et al.
Published: (2024)
Similar Items
-
Gromov-Wasserstein-like Distances in the Gaussian Mixture Models Space
by: Salmona, Antoine, et al.
Published: (2023) -
Min Generalized Sliced Gromov Wasserstein: A Scalable Path to Gromov Wasserstein
by: Shahbazi, Ashkan, et al.
Published: (2026) -
Shape-of-You: Fused Gromov-Wasserstein Optimal Transport for Semantic Correspondence in-the-Wild
by: Im, Jiin, et al.
Published: (2026) -
Improving Hyperbolic Representations via Gromov-Wasserstein Regularization
by: Yang, Yifei, et al.
Published: (2024) -
MOKD: Cross-domain Finetuning for Few-shot Classification via Maximizing Optimized Kernel Dependence
by: Tian, Hongduan, et al.
Published: (2024)