NeurIPS 2023 LLM Efficiency Fine-tuning Competition
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908272934518784 |
|---|---|
| author | Saroufim, Mark Perlitz, Yotam Choshen, Leshem Antiga, Luca Bowyer, Greg Puhrsch, Christian Guessous, Driss Rao, Supriya Chauhan, Geeta Kumar, Ashvini Kumar, Jindal Pawan Parikh, Rajpoot Ankur Isaacson, Joe Yang, Weiwei |
| author_facet | Saroufim, Mark Perlitz, Yotam Choshen, Leshem Antiga, Luca Bowyer, Greg Puhrsch, Christian Guessous, Driss Rao, Supriya Chauhan, Geeta Kumar, Ashvini Kumar, Jindal Pawan Parikh, Rajpoot Ankur Isaacson, Joe Yang, Weiwei |
| contents | Our analysis of the NeurIPS 2023 large language model (LLM) fine-tuning competition revealed the following trend: top-performing models exhibit significant overfitting on benchmark datasets, mirroring the broader issue of benchmark overfitting on popular leaderboards and that data curation is essential in order to get a high performing LLM. The competition, which consisted of two stages - an open evaluation stage with publicly available tasks and a closed evaluation stage with unseen tasks - allowed us to assess the generalizability of fine-tuned LLMs. Our results highlight the limitations of current benchmark-based evaluation schemes for generative models and demonstrate the need for more robust evaluation methods. Notably, the winning submissions utilized standard open-source libraries and focused primarily on data curation. To facilitate further research and promote reproducibility, we release all competition entries, Docker files, and evaluation infrastructure, providing a valuable resource for the community to explore fine-tuning, overfitting, and reproducibility in LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_13507 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | NeurIPS 2023 LLM Efficiency Fine-tuning Competition Saroufim, Mark Perlitz, Yotam Choshen, Leshem Antiga, Luca Bowyer, Greg Puhrsch, Christian Guessous, Driss Rao, Supriya Chauhan, Geeta Kumar, Ashvini Kumar, Jindal Pawan Parikh, Rajpoot Ankur Isaacson, Joe Yang, Weiwei Computation and Language Artificial Intelligence Our analysis of the NeurIPS 2023 large language model (LLM) fine-tuning competition revealed the following trend: top-performing models exhibit significant overfitting on benchmark datasets, mirroring the broader issue of benchmark overfitting on popular leaderboards and that data curation is essential in order to get a high performing LLM. The competition, which consisted of two stages - an open evaluation stage with publicly available tasks and a closed evaluation stage with unseen tasks - allowed us to assess the generalizability of fine-tuned LLMs. Our results highlight the limitations of current benchmark-based evaluation schemes for generative models and demonstrate the need for more robust evaluation methods. Notably, the winning submissions utilized standard open-source libraries and focused primarily on data curation. To facilitate further research and promote reproducibility, we release all competition entries, Docker files, and evaluation infrastructure, providing a valuable resource for the community to explore fine-tuning, overfitting, and reproducibility in LLMs. |
| title | NeurIPS 2023 LLM Efficiency Fine-tuning Competition |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2503.13507 |