How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Gordienko, Polina, Schollmeyer, Georg, Kreuter, Frauke, Jansen, Christoph |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking
by: Gordienko, Polina, et al.
Published: (2026)
by: Gordienko, Polina, et al.
Published: (2026)
Depth Functions for Partial Orders with a Descriptive Analysis of Machine Learning Algorithms
by: Blocher, Hannah, et al.
Published: (2023)
by: Blocher, Hannah, et al.
Published: (2023)
Comparing Machine Learning Algorithms by Union-Free Generic Depth
by: Blocher, Hannah, et al.
Published: (2023)
by: Blocher, Hannah, et al.
Published: (2023)
Reciprocal Learning
by: Rodemann, Julian, et al.
Published: (2024)
by: Rodemann, Julian, et al.
Published: (2024)
Robust Statistical Comparison of Random Variables with Locally Varying Scale of Measurement
by: Jansen, Christoph, et al.
Published: (2023)
by: Jansen, Christoph, et al.
Published: (2023)
Bias in the Loop: How Humans Evaluate AI-Generated Suggestions
by: Beck, Jacob, et al.
Published: (2025)
by: Beck, Jacob, et al.
Published: (2025)
Statistical Multicriteria Benchmarking via the GSD-Front
by: Jansen, Christoph, et al.
Published: (2024)
by: Jansen, Christoph, et al.
Published: (2024)
Robust Bayes Acts under Prior Perturbations: Contamination, Stability, and Selection Paths
by: Jansen, Christoph, et al.
Published: (2026)
by: Jansen, Christoph, et al.
Published: (2026)
Bias Begins with Data: The FairGround Corpus for Robust and Reproducible Research on Algorithmic Fairness
by: Simson, Jan, et al.
Published: (2025)
by: Simson, Jan, et al.
Published: (2025)
The Missing Link: Allocation Performance in Causal Machine Learning
by: Fischer-Abaigar, Unai, et al.
Published: (2024)
by: Fischer-Abaigar, Unai, et al.
Published: (2024)
Empirical Decision Theory
by: Jansen, Christoph, et al.
Published: (2025)
by: Jansen, Christoph, et al.
Published: (2025)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
by: Chew, Robert, et al.
Published: (2026)
by: Chew, Robert, et al.
Published: (2026)
Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction
by: Ball, Sarah, et al.
Published: (2025)
by: Ball, Sarah, et al.
Published: (2025)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
by: Ball, Sarah, et al.
Published: (2024)
by: Ball, Sarah, et al.
Published: (2024)
Bridging the gap: Towards an Expanded Toolkit for AI-driven Decision-Making in the Public Sector
by: Fischer-Abaigar, Unai, et al.
Published: (2023)
by: Fischer-Abaigar, Unai, et al.
Published: (2023)
Sources of Uncertainty in Supervised Machine Learning -- A Statisticians' View
by: Gruber, Cornelia, et al.
Published: (2023)
by: Gruber, Cornelia, et al.
Published: (2023)
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
by: Kern, Christoph, et al.
Published: (2023)
by: Kern, Christoph, et al.
Published: (2023)
Contributions to the Decision Theoretic Foundations of Machine Learning and Robust Statistics under Weakly Structured Information
by: Jansen, Christoph
Published: (2025)
by: Jansen, Christoph
Published: (2025)
LLM Robustness Leaderboard v1 --Technical report
by: Lefebvre, Pierre Peigné -, et al.
Published: (2025)
by: Lefebvre, Pierre Peigné -, et al.
Published: (2025)
Prompt-to-Leaderboard
by: Frick, Evan, et al.
Published: (2025)
by: Frick, Evan, et al.
Published: (2025)
Consensus in Motion: A Case of Dynamic Rationality of Sequential Learning in Probability Aggregation
by: Gordienko, Polina, et al.
Published: (2025)
by: Gordienko, Polina, et al.
Published: (2025)
On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
Measuring AI Progress in Drug Discovery: A Reproducible Leaderboard for the Tox21 Challenge
by: Ebner, Antonia, et al.
Published: (2025)
by: Ebner, Antonia, et al.
Published: (2025)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
by: Alzahrani, Norah, et al.
Published: (2024)
by: Alzahrani, Norah, et al.
Published: (2024)
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
by: Singh, Prabhant, et al.
Published: (2025)
by: Singh, Prabhant, et al.
Published: (2025)
A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation
by: Oyarhoseini, Hosna, et al.
Published: (2026)
by: Oyarhoseini, Hosna, et al.
Published: (2026)
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
by: Chen, Wenting, et al.
Published: (2025)
by: Chen, Wenting, et al.
Published: (2025)
How Hard is it to Confuse a World Model?
by: Radji, Waris, et al.
Published: (2025)
by: Radji, Waris, et al.
Published: (2025)
Make Still Further Progress: Chain of Thoughts for Tabular Data Leaderboard
by: Liu, Si-Yang, et al.
Published: (2025)
by: Liu, Si-Yang, et al.
Published: (2025)
DataDRILL: Formation Pressure Prediction and Kick Detection for Drilling Rigs
by: Arifeen, Murshedul, et al.
Published: (2024)
by: Arifeen, Murshedul, et al.
Published: (2024)
RepairBench: Leaderboard of Frontier Models for Program Repair
by: Silva, André, et al.
Published: (2024)
by: Silva, André, et al.
Published: (2024)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
by: Huang, Yangsibo, et al.
Published: (2025)
by: Huang, Yangsibo, et al.
Published: (2025)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
by: Suri, Anshuman, et al.
Published: (2025)
by: Suri, Anshuman, et al.
Published: (2025)
TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
by: Menezes, Michael, et al.
Published: (2025)
by: Menezes, Michael, et al.
Published: (2025)
How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning
by: Chen, Jiangwei, et al.
Published: (2026)
by: Chen, Jiangwei, et al.
Published: (2026)
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Audio2Rig: Artist-oriented deep learning tool for facial animation
by: Arcelin, Bastien, et al.
Published: (2024)
by: Arcelin, Bastien, et al.
Published: (2024)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
by: Topsakal, Oguzhan, et al.
Published: (2024)
by: Topsakal, Oguzhan, et al.
Published: (2024)
Identifying Models Behind Text-to-Image Leaderboards
by: Naseh, Ali, et al.
Published: (2026)
by: Naseh, Ali, et al.
Published: (2026)
Similar Items
-
Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking
by: Gordienko, Polina, et al.
Published: (2026) -
Depth Functions for Partial Orders with a Descriptive Analysis of Machine Learning Algorithms
by: Blocher, Hannah, et al.
Published: (2023) -
Comparing Machine Learning Algorithms by Union-Free Generic Depth
by: Blocher, Hannah, et al.
Published: (2023) -
Reciprocal Learning
by: Rodemann, Julian, et al.
Published: (2024) -
Robust Statistical Comparison of Random Variables with Locally Varying Scale of Measurement
by: Jansen, Christoph, et al.
Published: (2023)