LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Jiashuo, Wu, Jiayun, Wu, Chunjie, Liu, Jingkai, Wang, Zaiyuan, Zhou, Huan, Huang, Wenhao, Namkoong, Hongseok
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909975356375040
author Liu, Jiashuo
Wu, Jiayun
Wu, Chunjie
Liu, Jingkai
Wang, Zaiyuan
Zhou, Huan
Huang, Wenhao
Namkoong, Hongseok
author_facet Liu, Jiashuo
Wu, Jiayun
Wu, Chunjie
Liu, Jingkai
Wang, Zaiyuan
Zhou, Huan
Huang, Wenhao
Namkoong, Hongseok
contents The rapid proliferation of Large Language Models (LLMs) and diverse specialized benchmarks necessitates a shift from fragmented, task-specific metrics to a holistic, competitive ranking system that effectively aggregates performance across multiple ability dimensions. Primarily using static scoring, current evaluation methods are fundamentally limited. They struggle to determine the proper mix ratio across diverse benchmarks, and critically, they fail to capture a model's dynamic competitive fitness or its vulnerability when confronted with sequential, high-stakes tasks. To address this, we introduce the novel Competitive Swiss-System Dynamics (CSD) framework. CSD simulates a multi-round, sequential contest where models are dynamically paired across a curated sequence of benchmarks based on their accumulated win-loss record. And Monte Carlo Simulation ($N=100,000$ iterations) is used to approximate the statistically robust Expected Win Score ($E[S_m]$), which eliminates the noise of random pairing and early-round luck. Furthermore, we implement a Failure Sensitivity Analysis by parameterizing the per-round elimination quantity ($T_k$), which allows us to profile models based on their risk appetite--distinguishing between robust generalists and aggressive specialists. We demonstrate that CSD provides a more nuanced and context-aware ranking than traditional aggregate scoring and static pairwise models, representing a vital step towards risk-informed, next-generation LLM evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21010
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
Liu, Jiashuo
Wu, Jiayun
Wu, Chunjie
Liu, Jingkai
Wang, Zaiyuan
Zhou, Huan
Huang, Wenhao
Namkoong, Hongseok
Machine Learning
Artificial Intelligence
Performance
The rapid proliferation of Large Language Models (LLMs) and diverse specialized benchmarks necessitates a shift from fragmented, task-specific metrics to a holistic, competitive ranking system that effectively aggregates performance across multiple ability dimensions. Primarily using static scoring, current evaluation methods are fundamentally limited. They struggle to determine the proper mix ratio across diverse benchmarks, and critically, they fail to capture a model's dynamic competitive fitness or its vulnerability when confronted with sequential, high-stakes tasks. To address this, we introduce the novel Competitive Swiss-System Dynamics (CSD) framework. CSD simulates a multi-round, sequential contest where models are dynamically paired across a curated sequence of benchmarks based on their accumulated win-loss record. And Monte Carlo Simulation ($N=100,000$ iterations) is used to approximate the statistically robust Expected Win Score ($E[S_m]$), which eliminates the noise of random pairing and early-round luck. Furthermore, we implement a Failure Sensitivity Analysis by parameterizing the per-round elimination quantity ($T_k$), which allows us to profile models based on their risk appetite--distinguishing between robust generalists and aggressive specialists. We demonstrate that CSD provides a more nuanced and context-aware ranking than traditional aggregate scoring and static pairwise models, representing a vital step towards risk-informed, next-generation LLM evaluation.
title LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
topic Machine Learning
Artificial Intelligence
Performance
url https://arxiv.org/abs/2512.21010