The Pitfalls of Benchmarking in Algorithm Selection: What We Are Getting Wrong
Fuente:
arXiv
Guardado en:
| Autores principales: | Petelin, Gašper, Cenikj, Gjorgjina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Landscape Features in Single-Objective Continuous Optimization: Have We Hit a Wall in Algorithm Selection Generalization?
por: Cenikj, Gjorgjina, et al.
Publicado: (2025)
por: Cenikj, Gjorgjina, et al.
Publicado: (2025)
Comparing Optimization Algorithms Through the Lens of Search Behavior Analysis
por: Cenikj, Gjorgjina, et al.
Publicado: (2025)
por: Cenikj, Gjorgjina, et al.
Publicado: (2025)
ClustOpt: A Clustering-based Approach for Representing and Visualizing the Search Dynamics of Numerical Metaheuristic Optimization Algorithms
por: Cenikj, Gjorgjina, et al.
Publicado: (2025)
por: Cenikj, Gjorgjina, et al.
Publicado: (2025)
A Survey of Meta-features Used for Automated Selection of Algorithms for Black-box Single-objective Continuous Optimization
por: Cenikj, Gjorgjina, et al.
Publicado: (2024)
por: Cenikj, Gjorgjina, et al.
Publicado: (2024)
Evaluating Real-World Generalizability of Algorithm Selection Models
por: Cenikj, Gjorgjina, et al.
Publicado: (2026)
por: Cenikj, Gjorgjina, et al.
Publicado: (2026)
Instance Selection for Dynamic Algorithm Configuration with Reinforcement Learning: Improving Generalization
por: Benjamins, Carolin, et al.
Publicado: (2024)
por: Benjamins, Carolin, et al.
Publicado: (2024)
Generalization Ability of Feature-based Performance Prediction Models: A Statistical Analysis across Benchmarks
por: Nikolikj, Ana, et al.
Publicado: (2024)
por: Nikolikj, Ana, et al.
Publicado: (2024)
Starting Off on the Wrong Foot: Pitfalls in Data Preparation
por: Guo, Jiayi, et al.
Publicado: (2026)
por: Guo, Jiayi, et al.
Publicado: (2026)
Easy Problems That LLMs Get Wrong
por: Williams, Sean, et al.
Publicado: (2024)
por: Williams, Sean, et al.
Publicado: (2024)
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
por: Kjorvezir, Denica, et al.
Publicado: (2026)
por: Kjorvezir, Denica, et al.
Publicado: (2026)
What is Wrong with End-to-End Learning for Phase Retrieval?
por: Zhang, Wenjie, et al.
Publicado: (2024)
por: Zhang, Wenjie, et al.
Publicado: (2024)
What is Wrong with Perplexity for Long-context Language Modeling?
por: Fang, Lizhe, et al.
Publicado: (2024)
por: Fang, Lizhe, et al.
Publicado: (2024)
Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?
por: Wen, Xueru, et al.
Publicado: (2024)
por: Wen, Xueru, et al.
Publicado: (2024)
Overcoming Pitfalls in Graph Contrastive Learning Evaluation: Toward Comprehensive Benchmarks
por: Ma, Qian, et al.
Publicado: (2024)
por: Ma, Qian, et al.
Publicado: (2024)
Unsupervised Learning and Representation of Mandarin Tonal Categories by a Generative CNN
por: Schenck, Kai, et al.
Publicado: (2025)
por: Schenck, Kai, et al.
Publicado: (2025)
Don't Get Me Wrong: How to Apply Deep Visual Interpretations to Time Series
por: Loeffler, Christoffer, et al.
Publicado: (2022)
por: Loeffler, Christoffer, et al.
Publicado: (2022)
What's Wrong with Your Synthetic Tabular Data? Using Explainable AI to Evaluate Generative Models
por: Kapar, Jan, et al.
Publicado: (2025)
por: Kapar, Jan, et al.
Publicado: (2025)
Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability Detection
por: Risse, Niklas, et al.
Publicado: (2024)
por: Risse, Niklas, et al.
Publicado: (2024)
TabReD: Analyzing Pitfalls and Filling the Gaps in Tabular Deep Learning Benchmarks
por: Rubachev, Ivan, et al.
Publicado: (2024)
por: Rubachev, Ivan, et al.
Publicado: (2024)
Evaluating Supervised Machine Learning Models: Principles, Pitfalls, and Metric Selection
por: Liu, Xuanyan, et al.
Publicado: (2026)
por: Liu, Xuanyan, et al.
Publicado: (2026)
What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
por: Errica, Federico, et al.
Publicado: (2024)
por: Errica, Federico, et al.
Publicado: (2024)
We Need to Rethink Benchmarking in Anomaly Detection
por: Röchner, Philipp, et al.
Publicado: (2025)
por: Röchner, Philipp, et al.
Publicado: (2025)
The Hidden Pitfalls of the Cosine Similarity Loss
por: Draganov, Andrew, et al.
Publicado: (2024)
por: Draganov, Andrew, et al.
Publicado: (2024)
What You See is Not What You Get: Neural Partial Differential Equations and The Illusion of Learning
por: Mohan, Arvind, et al.
Publicado: (2024)
por: Mohan, Arvind, et al.
Publicado: (2024)
"Show Me What's Wrong!": Combining Charts and Text to Guide Data Analysis
por: Feliciano, Beatriz, et al.
Publicado: (2024)
por: Feliciano, Beatriz, et al.
Publicado: (2024)
What Can We Learn From MIMO Graph Convolutions?
por: Roth, Andreas, et al.
Publicado: (2025)
por: Roth, Andreas, et al.
Publicado: (2025)
On What We Can Learn from Low-Resolution Data
por: Frehr, Theresa Dahl, et al.
Publicado: (2026)
por: Frehr, Theresa Dahl, et al.
Publicado: (2026)
Response to Promises and Pitfalls of Deep Kernel Learning
por: Wilson, Andrew Gordon, et al.
Publicado: (2025)
por: Wilson, Andrew Gordon, et al.
Publicado: (2025)
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
por: Gekker, Gil, et al.
Publicado: (2025)
por: Gekker, Gil, et al.
Publicado: (2025)
What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data Slicing
por: Yang, Chenyang, et al.
Publicado: (2024)
por: Yang, Chenyang, et al.
Publicado: (2024)
The Pitfalls of KV Cache Compression
por: Chen, Alex, et al.
Publicado: (2025)
por: Chen, Alex, et al.
Publicado: (2025)
Out-of-Distribution Detection Methods Answer the Wrong Questions
por: Li, Yucen Lily, et al.
Publicado: (2025)
por: Li, Yucen Lily, et al.
Publicado: (2025)
Your Assumed DAG is Wrong and Here's How To Deal With It
por: Padh, Kirtan, et al.
Publicado: (2025)
por: Padh, Kirtan, et al.
Publicado: (2025)
Algorithm Selection with Probing Trajectories: Benchmarking the Choice of Classifier Model
por: Renau, Quentin, et al.
Publicado: (2025)
por: Renau, Quentin, et al.
Publicado: (2025)
Approaching an unknown communication system by latent space exploration and causal inference
por: Beguš, Gašper, et al.
Publicado: (2023)
por: Beguš, Gašper, et al.
Publicado: (2023)
Are We Winning the Wrong Game? Revisiting Evaluation Practices for Long-Term Time Series Forecasting
por: Phungtua-eng, Thanapol, et al.
Publicado: (2026)
por: Phungtua-eng, Thanapol, et al.
Publicado: (2026)
The Pitfalls and Promise of Conformal Inference Under Adversarial Attacks
por: Liu, Ziquan, et al.
Publicado: (2024)
por: Liu, Ziquan, et al.
Publicado: (2024)
What You See Is Not Always What You Get: Evaluating GPT's Comprehension of Source Code
por: Wen, Jiawen, et al.
Publicado: (2024)
por: Wen, Jiawen, et al.
Publicado: (2024)
Frugal Algorithm Selection
por: Kuş, Erdem, et al.
Publicado: (2024)
por: Kuş, Erdem, et al.
Publicado: (2024)
What is the Relationship between Tensor Factorizations and Circuits (and How Can We Exploit it)?
por: Loconte, Lorenzo, et al.
Publicado: (2024)
por: Loconte, Lorenzo, et al.
Publicado: (2024)
Ejemplares similares
-
Landscape Features in Single-Objective Continuous Optimization: Have We Hit a Wall in Algorithm Selection Generalization?
por: Cenikj, Gjorgjina, et al.
Publicado: (2025) -
Comparing Optimization Algorithms Through the Lens of Search Behavior Analysis
por: Cenikj, Gjorgjina, et al.
Publicado: (2025) -
ClustOpt: A Clustering-based Approach for Representing and Visualizing the Search Dynamics of Numerical Metaheuristic Optimization Algorithms
por: Cenikj, Gjorgjina, et al.
Publicado: (2025) -
A Survey of Meta-features Used for Automated Selection of Algorithms for Black-box Single-objective Continuous Optimization
por: Cenikj, Gjorgjina, et al.
Publicado: (2024) -
Evaluating Real-World Generalizability of Algorithm Selection Models
por: Cenikj, Gjorgjina, et al.
Publicado: (2026)