Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pareja, Aldo, Nayak, Nikhil Shivakumar, Wang, Hao, Killamsetty, Krishnateja, Sudalairaj, Shivchander, Zhao, Wenlong, Han, Seungwook, Bhandwaldar, Abhishek, Xu, Guangxuan, Xu, Kai, Han, Ligong, Inglis, Luke, Srivastava, Akash
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913616600498176
author Pareja, Aldo
Nayak, Nikhil Shivakumar
Wang, Hao
Killamsetty, Krishnateja
Sudalairaj, Shivchander
Zhao, Wenlong
Han, Seungwook
Bhandwaldar, Abhishek
Xu, Guangxuan
Xu, Kai
Han, Ligong
Inglis, Luke
Srivastava, Akash
author_facet Pareja, Aldo
Nayak, Nikhil Shivakumar
Wang, Hao
Killamsetty, Krishnateja
Sudalairaj, Shivchander
Zhao, Wenlong
Han, Seungwook
Bhandwaldar, Abhishek
Xu, Guangxuan
Xu, Kai
Han, Ligong
Inglis, Luke
Srivastava, Akash
contents The rise of large language models (LLMs) has created a significant disparity: industrial research labs with their computational resources, expert teams, and advanced infrastructures, can effectively fine-tune LLMs, while individual developers and small organizations face barriers due to limited resources. In this paper, we aim to bridge this gap by presenting a comprehensive study on supervised fine-tuning of LLMs using instruction-tuning datasets spanning diverse knowledge domains and skills. We focus on small-sized LLMs (3B to 7B parameters) for their cost-efficiency and accessibility. We explore various training configurations and strategies across four open-source pre-trained models. We provide detailed documentation of these configurations, revealing findings that challenge several common training practices, including hyperparameter recommendations from TULU and phased training recommended by Orca. Key insights from our work include: (i) larger batch sizes paired with lower learning rates lead to improved model performance on benchmarks such as MMLU, MTBench, and Open LLM Leaderboard; (ii) early-stage training dynamics, such as lower gradient norms and higher loss values, are strong indicators of better final model performance, enabling early termination of sub-optimal runs and significant computational savings; (iii) through a thorough exploration of hyperparameters like warmup steps and learning rate schedules, we provide guidance for practitioners and find that certain simplifications do not compromise performance; and (iv) we observed no significant difference in performance between phased and stacked training strategies, but stacked training is simpler and more sample efficient. With these findings holding robustly across datasets and models, we hope this study serves as a guide for practitioners fine-tuning small LLMs and promotes a more inclusive environment for LLM research.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13337
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
Pareja, Aldo
Nayak, Nikhil Shivakumar
Wang, Hao
Killamsetty, Krishnateja
Sudalairaj, Shivchander
Zhao, Wenlong
Han, Seungwook
Bhandwaldar, Abhishek
Xu, Guangxuan
Xu, Kai
Han, Ligong
Inglis, Luke
Srivastava, Akash
Machine Learning
Artificial Intelligence
53-04
I.2.7; I.2.6; I.2.4
The rise of large language models (LLMs) has created a significant disparity: industrial research labs with their computational resources, expert teams, and advanced infrastructures, can effectively fine-tune LLMs, while individual developers and small organizations face barriers due to limited resources. In this paper, we aim to bridge this gap by presenting a comprehensive study on supervised fine-tuning of LLMs using instruction-tuning datasets spanning diverse knowledge domains and skills. We focus on small-sized LLMs (3B to 7B parameters) for their cost-efficiency and accessibility. We explore various training configurations and strategies across four open-source pre-trained models. We provide detailed documentation of these configurations, revealing findings that challenge several common training practices, including hyperparameter recommendations from TULU and phased training recommended by Orca. Key insights from our work include: (i) larger batch sizes paired with lower learning rates lead to improved model performance on benchmarks such as MMLU, MTBench, and Open LLM Leaderboard; (ii) early-stage training dynamics, such as lower gradient norms and higher loss values, are strong indicators of better final model performance, enabling early termination of sub-optimal runs and significant computational savings; (iii) through a thorough exploration of hyperparameters like warmup steps and learning rate schedules, we provide guidance for practitioners and find that certain simplifications do not compromise performance; and (iv) we observed no significant difference in performance between phased and stacked training strategies, but stacked training is simpler and more sample efficient. With these findings holding robustly across datasets and models, we hope this study serves as a guide for practitioners fine-tuning small LLMs and promotes a more inclusive environment for LLM research.
title Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
topic Machine Learning
Artificial Intelligence
53-04
I.2.7; I.2.6; I.2.4
url https://arxiv.org/abs/2412.13337