How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Prabhant, Hess, Sibylle, Vanschoren, Joaquin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Occam's model: Selecting simpler representations for better transferability estimation
by: Singh, Prabhant, et al.
Published: (2025)
by: Singh, Prabhant, et al.
Published: (2025)
Meta-Learning for Unsupervised Outlier Detection with Optimal Transport
by: Singh, Prabhant, et al.
Published: (2022)
by: Singh, Prabhant, et al.
Published: (2022)
On Supernet Transfer Learning for Effective Task Adaptation
by: Singh, Prabhant, et al.
Published: (2024)
by: Singh, Prabhant, et al.
Published: (2024)
CLAMS: A System for Zero-Shot Model Selection for Clustering
by: Singh, Prabhant, et al.
Published: (2024)
by: Singh, Prabhant, et al.
Published: (2024)
Meta-Learning Transformers to Improve In-Context Generalization
by: Braccaioli, Lorenzo, et al.
Published: (2025)
by: Braccaioli, Lorenzo, et al.
Published: (2025)
Robustness of AutoML on Dirty Categorical Data
by: Bueno, Marcos L. P., et al.
Published: (2026)
by: Bueno, Marcos L. P., et al.
Published: (2026)
The Evaluation Game: Beyond Static LLM Benchmarking
by: Wang, Paul, et al.
Published: (2026)
by: Wang, Paul, et al.
Published: (2026)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
Automatic Combination of Sample Selection Strategies for Few-Shot Learning
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
Beyond a Single Perspective: Towards a Realistic Evaluation of Website Fingerprinting Attacks
by: Deng, Xinhao, et al.
Published: (2025)
by: Deng, Xinhao, et al.
Published: (2025)
Can time series forecasting be automated? A benchmark and analysis
by: Sreedhara, Anvitha Thirthapura, et al.
Published: (2024)
by: Sreedhara, Anvitha Thirthapura, et al.
Published: (2024)
LLM Robustness Leaderboard v1 --Technical report
by: Lefebvre, Pierre Peigné -, et al.
Published: (2025)
by: Lefebvre, Pierre Peigné -, et al.
Published: (2025)
How predictable is language model benchmark performance?
by: Owen, David
Published: (2024)
by: Owen, David
Published: (2024)
Automated Machine Learning for Unsupervised Tabular Tasks
by: Singh, Prabhant, et al.
Published: (2025)
by: Singh, Prabhant, et al.
Published: (2025)
Automated Reinforcement Learning: An Overview
by: Afshar, Reza Refaei, et al.
Published: (2022)
by: Afshar, Reza Refaei, et al.
Published: (2022)
FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning
by: Hu, Liang, et al.
Published: (2025)
by: Hu, Liang, et al.
Published: (2025)
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
by: Chen, Wenting, et al.
Published: (2025)
by: Chen, Wenting, et al.
Published: (2025)
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
Towards impactful challenges: post-challenge paper, benchmarks and other dissemination actions
by: Marot, Antoine, et al.
Published: (2023)
by: Marot, Antoine, et al.
Published: (2023)
MetaLead: A Comprehensive Human-Curated Leaderboard Dataset for Transparent Reporting of Machine Learning Experiments
by: Timmer, Roelien C., et al.
Published: (2026)
by: Timmer, Roelien C., et al.
Published: (2026)
EXACT: Towards a platform for empirically benchmarking Machine Learning model explanation methods
by: Clark, Benedict, et al.
Published: (2024)
by: Clark, Benedict, et al.
Published: (2024)
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
by: Alzahrani, Norah, et al.
Published: (2024)
by: Alzahrani, Norah, et al.
Published: (2024)
Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments
by: Delgrange, Florent
Published: (2026)
by: Delgrange, Florent
Published: (2026)
Beyond Static Uncertainty: Modeling Temporal Uncertainty Dynamics for Probabilistic Time Series Forecasting
by: Wang, Yijun, et al.
Published: (2026)
by: Wang, Yijun, et al.
Published: (2026)
MAWIFlow Benchmark: Realistic Flow-Based Evaluation for Network Intrusion Detection
by: Schraven, Joshua, et al.
Published: (2025)
by: Schraven, Joshua, et al.
Published: (2025)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
by: Chanin, David, et al.
Published: (2026)
by: Chanin, David, et al.
Published: (2026)
How Deep is your Guess? A Fresh Perspective on Deep Learning for Medical Time-Series Imputation
by: Qian, Linglong, et al.
Published: (2024)
by: Qian, Linglong, et al.
Published: (2024)
Continual Pre-training of MoEs: How robust is your router?
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
The role of positional encodings in the ARC benchmark
by: Costa, Guilherme H. Bandeira, et al.
Published: (2025)
by: Costa, Guilherme H. Bandeira, et al.
Published: (2025)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
by: Rao, Varun, et al.
Published: (2025)
by: Rao, Varun, et al.
Published: (2025)
The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next
by: Amin, Adil
Published: (2026)
by: Amin, Adil
Published: (2026)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
by: Brack, Manuel, et al.
Published: (2025)
by: Brack, Manuel, et al.
Published: (2025)
Secret mixtures of experts inside your LLM
by: Boix-Adsera, Enric
Published: (2025)
by: Boix-Adsera, Enric
Published: (2025)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
by: Kinniment, Megan, et al.
Published: (2023)
by: Kinniment, Megan, et al.
Published: (2023)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
by: Topsakal, Oguzhan, et al.
Published: (2024)
by: Topsakal, Oguzhan, et al.
Published: (2024)
Monitoring of Static Fairness
by: Henzinger, Thomas A., et al.
Published: (2025)
by: Henzinger, Thomas A., et al.
Published: (2025)
GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
Toward industrial use of continual learning : new metrics proposal for class incremental learning
by: Abbas, Konaté Mohamed, et al.
Published: (2024)
by: Abbas, Konaté Mohamed, et al.
Published: (2024)
DAVE: Diagnostic benchmark for Audio Visual Evaluation
by: Radevski, Gorjan, et al.
Published: (2025)
by: Radevski, Gorjan, et al.
Published: (2025)
A More Realistic Evaluation of Cross-Frequency Transfer Learning and Foundation Forecasting Models
by: Olivares, Kin G., et al.
Published: (2025)
by: Olivares, Kin G., et al.
Published: (2025)
Similar Items
-
Occam's model: Selecting simpler representations for better transferability estimation
by: Singh, Prabhant, et al.
Published: (2025) -
Meta-Learning for Unsupervised Outlier Detection with Optimal Transport
by: Singh, Prabhant, et al.
Published: (2022) -
On Supernet Transfer Learning for Effective Task Adaptation
by: Singh, Prabhant, et al.
Published: (2024) -
CLAMS: A System for Zero-Shot Model Selection for Clustering
by: Singh, Prabhant, et al.
Published: (2024) -
Meta-Learning Transformers to Improve In-Context Generalization
by: Braccaioli, Lorenzo, et al.
Published: (2025)