NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yagoubi, Mouadh, Dahou, Yasser, Mokeddem, Billel, Belkada, Younes, Le-Khac, Phuc H., Boussaha, Basma El Amel, Alami, Reda, Zuo, Jingwei, Marsili, Damiano, Farooq, Mugariya, Lalmas, Mounia, Gkioxari, Georgia, Gallinari, Patrick, Torr, Philip, Hacid, Hakim
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915333521014784
author Yagoubi, Mouadh
Dahou, Yasser
Mokeddem, Billel
Belkada, Younes
Le-Khac, Phuc H.
Boussaha, Basma El Amel
Alami, Reda
Zuo, Jingwei
Marsili, Damiano
Farooq, Mugariya
Lalmas, Mounia
Gkioxari, Georgia
Gallinari, Patrick
Torr, Philip
Hacid, Hakim
author_facet Yagoubi, Mouadh
Dahou, Yasser
Mokeddem, Billel
Belkada, Younes
Le-Khac, Phuc H.
Boussaha, Basma El Amel
Alami, Reda
Zuo, Jingwei
Marsili, Damiano
Farooq, Mugariya
Lalmas, Mounia
Gkioxari, Georgia
Gallinari, Patrick
Torr, Philip
Hacid, Hakim
contents Existing benchmarks have proven effective for assessing the performance of fully trained large language models. However, we find striking differences in the early training stages of small models, where benchmarks often fail to provide meaningful or discriminative signals. To explore how these differences arise, this competition tackles the challenge of designing scientific knowledge evaluation tasks specifically tailored for measuring early training progress of language models. Participants are invited to develop novel evaluation methodologies or adapt existing benchmarks to better capture performance differences among language models. To support this effort, we provide three pre-trained small models (0.5B, 1B, and 3B parameters), along with intermediate checkpoints sampled during training up to 200B tokens. All experiments and development work can be run on widely available free cloud-based GPU platforms, making participation accessible to researchers with limited computational resources. Submissions will be evaluated based on three criteria: the quality of the performance signal they produce, the consistency of model rankings at 1 trillion tokens of training, and their relevance to the scientific knowledge domain. By promoting the design of tailored evaluation strategies for early training, this competition aims to attract a broad range of participants from various disciplines, including those who may not be machine learning experts or have access to dedicated GPU resources. Ultimately, this initiative seeks to make foundational LLM research more systematic and benchmark-informed from the earliest phases of model development.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07731
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
Yagoubi, Mouadh
Dahou, Yasser
Mokeddem, Billel
Belkada, Younes
Le-Khac, Phuc H.
Boussaha, Basma El Amel
Alami, Reda
Zuo, Jingwei
Marsili, Damiano
Farooq, Mugariya
Lalmas, Mounia
Gkioxari, Georgia
Gallinari, Patrick
Torr, Philip
Hacid, Hakim
Artificial Intelligence
Existing benchmarks have proven effective for assessing the performance of fully trained large language models. However, we find striking differences in the early training stages of small models, where benchmarks often fail to provide meaningful or discriminative signals. To explore how these differences arise, this competition tackles the challenge of designing scientific knowledge evaluation tasks specifically tailored for measuring early training progress of language models. Participants are invited to develop novel evaluation methodologies or adapt existing benchmarks to better capture performance differences among language models. To support this effort, we provide three pre-trained small models (0.5B, 1B, and 3B parameters), along with intermediate checkpoints sampled during training up to 200B tokens. All experiments and development work can be run on widely available free cloud-based GPU platforms, making participation accessible to researchers with limited computational resources. Submissions will be evaluated based on three criteria: the quality of the performance signal they produce, the consistency of model rankings at 1 trillion tokens of training, and their relevance to the scientific knowledge domain. By promoting the design of tailored evaluation strategies for early training, this competition aims to attract a broad range of participants from various disciplines, including those who may not be machine learning experts or have access to dedicated GPU resources. Ultimately, this initiative seeks to make foundational LLM research more systematic and benchmark-informed from the earliest phases of model development.
title NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2506.07731