Selecting Large Language Model to Fine-tune via Rectified Scaling Law

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Haowei, Huang, Baizhou, Ye, Haotian, Chen, Qinyu, Wang, Zihao, Li, Sujian, Ma, Jianzhu, Wan, Xiaojun, Zou, James, Liang, Yitao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910460793585664
author Lin, Haowei
Huang, Baizhou
Ye, Haotian
Chen, Qinyu
Wang, Zihao
Li, Sujian
Ma, Jianzhu
Wan, Xiaojun
Zou, James
Liang, Yitao
author_facet Lin, Haowei
Huang, Baizhou
Ye, Haotian
Chen, Qinyu
Wang, Zihao
Li, Sujian
Ma, Jianzhu
Wan, Xiaojun
Zou, James
Liang, Yitao
contents The ever-growing ecosystem of LLMs has posed a challenge in selecting the most appropriate pre-trained model to fine-tune amidst a sea of options. Given constrained resources, fine-tuning all models and making selections afterward is unrealistic. In this work, we formulate this resource-constrained selection task into predicting fine-tuning performance and illustrate its natural connection with Scaling Law. Unlike pre-training, we find that the fine-tuning scaling curve includes not just the well-known "power phase" but also the previously unobserved "pre-power phase". We also explain why existing Scaling Law fails to capture this phase transition phenomenon both theoretically and empirically. To address this, we introduce the concept of "pre-learned data size" into our Rectified Scaling Law, which overcomes theoretical limitations and fits experimental results much better. By leveraging our law, we propose a novel LLM selection algorithm that selects the near-optimal model with hundreds of times less resource consumption, while other methods may provide negatively correlated selection. The project page is available at rectified-scaling-law.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02314
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Selecting Large Language Model to Fine-tune via Rectified Scaling Law
Lin, Haowei
Huang, Baizhou
Ye, Haotian
Chen, Qinyu
Wang, Zihao
Li, Sujian
Ma, Jianzhu
Wan, Xiaojun
Zou, James
Liang, Yitao
Machine Learning
Artificial Intelligence
Computation and Language
The ever-growing ecosystem of LLMs has posed a challenge in selecting the most appropriate pre-trained model to fine-tune amidst a sea of options. Given constrained resources, fine-tuning all models and making selections afterward is unrealistic. In this work, we formulate this resource-constrained selection task into predicting fine-tuning performance and illustrate its natural connection with Scaling Law. Unlike pre-training, we find that the fine-tuning scaling curve includes not just the well-known "power phase" but also the previously unobserved "pre-power phase". We also explain why existing Scaling Law fails to capture this phase transition phenomenon both theoretically and empirically. To address this, we introduce the concept of "pre-learned data size" into our Rectified Scaling Law, which overcomes theoretical limitations and fits experimental results much better. By leveraging our law, we propose a novel LLM selection algorithm that selects the near-optimal model with hundreds of times less resource consumption, while other methods may provide negatively correlated selection. The project page is available at rectified-scaling-law.github.io.
title Selecting Large Language Model to Fine-tune via Rectified Scaling Law
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2402.02314