ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Han, Li, Kehan, Li, Dongbai, He, Yue, Zhang, Xingxuan, Cui, Peng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915588470734848
author Yu, Han
Li, Kehan
Li, Dongbai
He, Yue
Zhang, Xingxuan
Cui, Peng
author_facet Yu, Han
Li, Kehan
Li, Dongbai
He, Yue
Zhang, Xingxuan
Cui, Peng
contents Recently, there has been gradually more attention paid to Out-of-Distribution (OOD) performance prediction, whose goal is to predict the performance of trained models on unlabeled OOD test datasets, so that we could better leverage and deploy off-the-shelf trained models in risk-sensitive scenarios. Although progress has been made in this area, evaluation protocols in previous literature are inconsistent, and most works cover only a limited number of real-world OOD datasets and types of distribution shifts. To provide convenient and fair comparisons for various algorithms, we propose Out-of-Distribution Performance Prediction Benchmark (ODP-Bench), a comprehensive benchmark that includes most commonly used OOD datasets and existing practical performance prediction algorithms. We provide our trained models as a testbench for future researchers, thus guaranteeing the consistency of comparison and avoiding the burden of repeating the model training process. Furthermore, we also conduct in-depth experimental analyses to better understand their capability boundary.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27263
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
Yu, Han
Li, Kehan
Li, Dongbai
He, Yue
Zhang, Xingxuan
Cui, Peng
Machine Learning
Recently, there has been gradually more attention paid to Out-of-Distribution (OOD) performance prediction, whose goal is to predict the performance of trained models on unlabeled OOD test datasets, so that we could better leverage and deploy off-the-shelf trained models in risk-sensitive scenarios. Although progress has been made in this area, evaluation protocols in previous literature are inconsistent, and most works cover only a limited number of real-world OOD datasets and types of distribution shifts. To provide convenient and fair comparisons for various algorithms, we propose Out-of-Distribution Performance Prediction Benchmark (ODP-Bench), a comprehensive benchmark that includes most commonly used OOD datasets and existing practical performance prediction algorithms. We provide our trained models as a testbench for future researchers, thus guaranteeing the consistency of comparison and avoiding the burden of repeating the model training process. Furthermore, we also conduct in-depth experimental analyses to better understand their capability boundary.
title ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
topic Machine Learning
url https://arxiv.org/abs/2510.27263