Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yue, Yuanhao, Wang, Chengyu, Huang, Jun, Wang, Peng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908157788291072
author Yue, Yuanhao
Wang, Chengyu
Huang, Jun
Wang, Peng
author_facet Yue, Yuanhao
Wang, Chengyu
Huang, Jun
Wang, Peng
contents Specializing LLMs in various domain-specific tasks has emerged as a critical step towards achieving high performance. However, the construction and annotation of datasets in specific domains are always very costly. Apart from using superior and expensive closed-source LLM APIs to construct datasets, some open-source models have become strong enough to handle dataset construction in many scenarios. Thus, we present a family of data augmentation models designed to significantly improve the efficiency for model fine-tuning. These models, trained based on sufficiently small LLMs, support key functionalities with low inference costs: instruction expansion, instruction refinement, and instruction-response pair expansion. To fulfill this goal, we first construct an automatic data collection system with seed datasets generated from both public repositories and our in-house datasets. This system leverages powerful LLMs to expand, refine and re-write the instructions and responses, incorporating quality assessment techniques. Following this, we introduce the training process of our models, which effectively distills task-solving and text synthesis abilities from teacher LLMs. Finally, we demonstrate how we integrate these functionalities into a machine learning platform to support low-cost LLM fine-tuning from both dataset preparation and training perspectives for users. Experiments and an application study prove the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04871
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud
Yue, Yuanhao
Wang, Chengyu
Huang, Jun
Wang, Peng
Computation and Language
Specializing LLMs in various domain-specific tasks has emerged as a critical step towards achieving high performance. However, the construction and annotation of datasets in specific domains are always very costly. Apart from using superior and expensive closed-source LLM APIs to construct datasets, some open-source models have become strong enough to handle dataset construction in many scenarios. Thus, we present a family of data augmentation models designed to significantly improve the efficiency for model fine-tuning. These models, trained based on sufficiently small LLMs, support key functionalities with low inference costs: instruction expansion, instruction refinement, and instruction-response pair expansion. To fulfill this goal, we first construct an automatic data collection system with seed datasets generated from both public repositories and our in-house datasets. This system leverages powerful LLMs to expand, refine and re-write the instructions and responses, incorporating quality assessment techniques. Following this, we introduce the training process of our models, which effectively distills task-solving and text synthesis abilities from teacher LLMs. Finally, we demonstrate how we integrate these functionalities into a machine learning platform to support low-cost LLM fine-tuning from both dataset preparation and training perspectives for users. Experiments and an application study prove the effectiveness of our approach.
title Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud
topic Computation and Language
url https://arxiv.org/abs/2412.04871