Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Belcak, Peter, Heinrich, Greg, Kautz, Jan, Molchanov, Pavlo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916800177897472
author Belcak, Peter
Heinrich, Greg
Kautz, Jan
Molchanov, Pavlo
author_facet Belcak, Peter
Heinrich, Greg
Kautz, Jan
Molchanov, Pavlo
contents Finetuning language models for a new domain inevitably leads to the deterioration of their general performance. This becomes more pronounced the more limited the finetuning data resource. We introduce minifinetuning (MFT), a method for language model domain adaptation that considerably reduces the effects of overfitting-induced degeneralization in low-data settings and which does so in the absence of any pre-training data for replay. MFT demonstrates 2-10x more favourable specialization-to-degeneralization ratios than standard finetuning across a wide range of models and domains and exhibits an intrinsic robustness to overfitting when data in the new domain is scarce and down to as little as 500 samples. Employing corrective self-distillation that is individualized on the sample level, MFT outperforms parameter-efficient finetuning methods, demonstrates replay-like degeneralization mitigation properties, and is composable with either for a combined effect.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15702
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation
Belcak, Peter
Heinrich, Greg
Kautz, Jan
Molchanov, Pavlo
Machine Learning
Artificial Intelligence
Finetuning language models for a new domain inevitably leads to the deterioration of their general performance. This becomes more pronounced the more limited the finetuning data resource. We introduce minifinetuning (MFT), a method for language model domain adaptation that considerably reduces the effects of overfitting-induced degeneralization in low-data settings and which does so in the absence of any pre-training data for replay. MFT demonstrates 2-10x more favourable specialization-to-degeneralization ratios than standard finetuning across a wide range of models and domains and exhibits an intrinsic robustness to overfitting when data in the new domain is scarce and down to as little as 500 samples. Employing corrective self-distillation that is individualized on the sample level, MFT outperforms parameter-efficient finetuning methods, demonstrates replay-like degeneralization mitigation properties, and is composable with either for a combined effect.
title Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.15702