NeurIPS 2023 LLM Efficiency Fine-tuning Competition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Saroufim, Mark, Perlitz, Yotam, Choshen, Leshem, Antiga, Luca, Bowyer, Greg, Puhrsch, Christian, Guessous, Driss, Rao, Supriya, Chauhan, Geeta, Kumar, Ashvini, Kumar, Jindal Pawan, Parikh, Rajpoot Ankur, Isaacson, Joe, Yang, Weiwei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908272934518784
author Saroufim, Mark
Perlitz, Yotam
Choshen, Leshem
Antiga, Luca
Bowyer, Greg
Puhrsch, Christian
Guessous, Driss
Rao, Supriya
Chauhan, Geeta
Kumar, Ashvini
Kumar, Jindal Pawan
Parikh, Rajpoot Ankur
Isaacson, Joe
Yang, Weiwei
author_facet Saroufim, Mark
Perlitz, Yotam
Choshen, Leshem
Antiga, Luca
Bowyer, Greg
Puhrsch, Christian
Guessous, Driss
Rao, Supriya
Chauhan, Geeta
Kumar, Ashvini
Kumar, Jindal Pawan
Parikh, Rajpoot Ankur
Isaacson, Joe
Yang, Weiwei
contents Our analysis of the NeurIPS 2023 large language model (LLM) fine-tuning competition revealed the following trend: top-performing models exhibit significant overfitting on benchmark datasets, mirroring the broader issue of benchmark overfitting on popular leaderboards and that data curation is essential in order to get a high performing LLM. The competition, which consisted of two stages - an open evaluation stage with publicly available tasks and a closed evaluation stage with unseen tasks - allowed us to assess the generalizability of fine-tuned LLMs. Our results highlight the limitations of current benchmark-based evaluation schemes for generative models and demonstrate the need for more robust evaluation methods. Notably, the winning submissions utilized standard open-source libraries and focused primarily on data curation. To facilitate further research and promote reproducibility, we release all competition entries, Docker files, and evaluation infrastructure, providing a valuable resource for the community to explore fine-tuning, overfitting, and reproducibility in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NeurIPS 2023 LLM Efficiency Fine-tuning Competition
Saroufim, Mark
Perlitz, Yotam
Choshen, Leshem
Antiga, Luca
Bowyer, Greg
Puhrsch, Christian
Guessous, Driss
Rao, Supriya
Chauhan, Geeta
Kumar, Ashvini
Kumar, Jindal Pawan
Parikh, Rajpoot Ankur
Isaacson, Joe
Yang, Weiwei
Computation and Language
Artificial Intelligence
Our analysis of the NeurIPS 2023 large language model (LLM) fine-tuning competition revealed the following trend: top-performing models exhibit significant overfitting on benchmark datasets, mirroring the broader issue of benchmark overfitting on popular leaderboards and that data curation is essential in order to get a high performing LLM. The competition, which consisted of two stages - an open evaluation stage with publicly available tasks and a closed evaluation stage with unseen tasks - allowed us to assess the generalizability of fine-tuned LLMs. Our results highlight the limitations of current benchmark-based evaluation schemes for generative models and demonstrate the need for more robust evaluation methods. Notably, the winning submissions utilized standard open-source libraries and focused primarily on data curation. To facilitate further research and promote reproducibility, we release all competition entries, Docker files, and evaluation infrastructure, providing a valuable resource for the community to explore fine-tuning, overfitting, and reproducibility in LLMs.
title NeurIPS 2023 LLM Efficiency Fine-tuning Competition
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.13507