Systematische Untersuchung: Finetuning von Sprachmodellen

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Reetz
Natura: Recurso digital
Lingua:tedesco
Pubblicazione: Zenodo 2024
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866901527480762368
author Reetz
author_facet Reetz
contents <p><strong>The entire work is written in german, but here is a summary translated in english.</strong></p> <p><strong>----</strong></p> <p>The subject of this work is the systematic investigation of fine-tuning language models.<br>In this study, four models were examined: Llama2-7B-hf, microsoft-Phi-3-medium-4k-instruct, tiiuae-falcon-7b, and openai-community-gpt2. The models were quantized and fine-tuned with LoRA, utilizing various techniques. The aim of the investigation was to analyze the performance and the quality of responses on both general and domain-specific datasets, depending on the number of training epochs, and to gain insights into the fine-tuning process.</p> <p>Fine-tuning with LoRA shows that a higher LoRA rank and longer training can lead to the model forgetting (practically overwriting) more and faster of its pre-trained knowledge, while a smaller rank allows for better generalization. A disadvantage of this strategy is that the model learns more slowly.<br>Since the focus was ultimately on the technical implementation and the details, the original goal was somewhat overlooked. Fine-tuning small language models does not necessarily require a large dataset or many epochs. However, it is important to review the choice of hyperparameters; the results, however, were only evaluated at the end of this work, leaving little room for adjustments.</p> <p>Moreover, the use of additional techniques and methods, such as FlashAttention, could increase the model's efficiency in terms of training and inference. Quantizing models and loading large models layer by layer allows them to be run on consumer-grade GPUs. However, this comes with an increased latency in inference and a degradation in model quality.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_15684931
institution Zenodo
language deu
publishDate 2024
publisher Zenodo
record_format zenodo
spellingShingle Systematische Untersuchung: Finetuning von Sprachmodellen
Reetz
<p><strong>The entire work is written in german, but here is a summary translated in english.</strong></p> <p><strong>----</strong></p> <p>The subject of this work is the systematic investigation of fine-tuning language models.<br>In this study, four models were examined: Llama2-7B-hf, microsoft-Phi-3-medium-4k-instruct, tiiuae-falcon-7b, and openai-community-gpt2. The models were quantized and fine-tuned with LoRA, utilizing various techniques. The aim of the investigation was to analyze the performance and the quality of responses on both general and domain-specific datasets, depending on the number of training epochs, and to gain insights into the fine-tuning process.</p> <p>Fine-tuning with LoRA shows that a higher LoRA rank and longer training can lead to the model forgetting (practically overwriting) more and faster of its pre-trained knowledge, while a smaller rank allows for better generalization. A disadvantage of this strategy is that the model learns more slowly.<br>Since the focus was ultimately on the technical implementation and the details, the original goal was somewhat overlooked. Fine-tuning small language models does not necessarily require a large dataset or many epochs. However, it is important to review the choice of hyperparameters; the results, however, were only evaluated at the end of this work, leaving little room for adjustments.</p> <p>Moreover, the use of additional techniques and methods, such as FlashAttention, could increase the model's efficiency in terms of training and inference. Quantizing models and loading large models layer by layer allows them to be run on consumer-grade GPUs. However, this comes with an increased latency in inference and a degradation in model quality.</p>
title Systematische Untersuchung: Finetuning von Sprachmodellen
url https://doi.org/10.5281/zenodo.15684931