LLM Olympiad: Why Model Evaluation Needs a Sealed Exam

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cruz, Jan Christian Blaise, Aji, Alham Fikri
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910069142061056
author Cruz, Jan Christian Blaise
Aji, Alham Fikri
author_facet Cruz, Jan Christian Blaise
Aji, Alham Fikri
contents Benchmarks and leaderboards are how NLP most often communicates progress, but in the LLM era they are increasingly easy to misread. Scores can reflect benchmark-chasing, hidden evaluation choices, or accidental exposure to test content -- not just broad capability. Closed benchmarks delay some of these issues, but reduce transparency and make it harder for the community to learn from results. We argue for a complementary practice: an Olympiad-style evaluation event where problems are sealed until evaluation, submissions are frozen in advance, and all entries run through one standardized harness. After scoring, the full task set and evaluation code are released so results can be reproduced and audited. This design aims to make strong performance harder to ``manufacture'' and easier to trust.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23292
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
Cruz, Jan Christian Blaise
Aji, Alham Fikri
Artificial Intelligence
Computation and Language
Benchmarks and leaderboards are how NLP most often communicates progress, but in the LLM era they are increasingly easy to misread. Scores can reflect benchmark-chasing, hidden evaluation choices, or accidental exposure to test content -- not just broad capability. Closed benchmarks delay some of these issues, but reduce transparency and make it harder for the community to learn from results. We argue for a complementary practice: an Olympiad-style evaluation event where problems are sealed until evaluation, submissions are frozen in advance, and all entries run through one standardized harness. After scoring, the full task set and evaluation code are released so results can be reproduced and audited. This design aims to make strong performance harder to ``manufacture'' and easier to trust.
title LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.23292