Saved in:
Bibliographic Details
Main Authors: Cai, Rui, Mo, Weijie Jacky, Wen, Xiaofei, Ma, Qiyao, Zhu, Wenhui, Chen, Xiwen, Chen, Muhao, Zhao, Zhe
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.07075
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917471119736832
author Cai, Rui
Mo, Weijie Jacky
Wen, Xiaofei
Ma, Qiyao
Zhu, Wenhui
Chen, Xiwen
Chen, Muhao
Zhao, Zhe
author_facet Cai, Rui
Mo, Weijie Jacky
Wen, Xiaofei
Ma, Qiyao
Zhu, Wenhui
Chen, Xiwen
Chen, Muhao
Zhao, Zhe
contents The open-source model ecosystem now contains hundreds of thousands of pretrained models, yet picking the best model for a new dataset is increasingly infeasible: new models and unbenchmarked datasets emerge continuously, leaving practitioners with no prior records on either side. Existing approaches handle only fragments of this in-the-wild setting: AutoML and transferability estimation select models from small predefined pools or require expensive per-model forward passes on the target dataset, while model routing presupposes a given candidate pool. We introduce ModelLens, a unified framework for model recommendation in the wild. Our key insight is that public leaderboard interactions, though scattered and noisy, collectively trace out an implicit atlas of model capabilities across heterogeneous evaluation settings, a signal rich enough to learn from directly. By learning a performance-aware latent space over model--dataset--metric tuples, ModelLens ranks unseen models on unseen datasets without running candidates on the target dataset. On a new benchmark of 1.62M evaluation records spanning 47K models and 9.6K datasets, ModelLens surpasses baselines that either rely on metadata alone or require running each candidate on the target dataset. Its recommended Top-K pools further improve multiple representative routing methods by up to 81% across diverse QA benchmarks. Case studies on recently released benchmarks further confirm generalization to both text and vision-language tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07075
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ModelLens: Finding the Best for Your Task from Myriads of Models
Cai, Rui
Mo, Weijie Jacky
Wen, Xiaofei
Ma, Qiyao
Zhu, Wenhui
Chen, Xiwen
Chen, Muhao
Zhao, Zhe
Machine Learning
The open-source model ecosystem now contains hundreds of thousands of pretrained models, yet picking the best model for a new dataset is increasingly infeasible: new models and unbenchmarked datasets emerge continuously, leaving practitioners with no prior records on either side. Existing approaches handle only fragments of this in-the-wild setting: AutoML and transferability estimation select models from small predefined pools or require expensive per-model forward passes on the target dataset, while model routing presupposes a given candidate pool. We introduce ModelLens, a unified framework for model recommendation in the wild. Our key insight is that public leaderboard interactions, though scattered and noisy, collectively trace out an implicit atlas of model capabilities across heterogeneous evaluation settings, a signal rich enough to learn from directly. By learning a performance-aware latent space over model--dataset--metric tuples, ModelLens ranks unseen models on unseen datasets without running candidates on the target dataset. On a new benchmark of 1.62M evaluation records spanning 47K models and 9.6K datasets, ModelLens surpasses baselines that either rely on metadata alone or require running each candidate on the target dataset. Its recommended Top-K pools further improve multiple representative routing methods by up to 81% across diverse QA benchmarks. Case studies on recently released benchmarks further confirm generalization to both text and vision-language tasks.
title ModelLens: Finding the Best for Your Task from Myriads of Models
topic Machine Learning
url https://arxiv.org/abs/2605.07075