DoGE: Domain Reweighting with Generalization Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Simin, Pagliardini, Matteo, Jaggi, Martin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909092098867200
author Fan, Simin
Pagliardini, Matteo
Jaggi, Martin
author_facet Fan, Simin
Pagliardini, Matteo
Jaggi, Martin
contents The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs). Despite its importance, recent LLMs still rely on heuristics and trial and error to increase or reduce the influence of data-domains. We propose DOmain reweighting with Generalization Estimation (DoGE), which optimizes the probability of sampling from each domain (domain weights) in a principled way. Our approach is a two-stage process consisting of (i) training a proxy model to obtain domain weights using a bi-level optimization algorithm; (ii) training a larger base model by sampling training domains according to the learned domain weights. In our experiments, we extensively show how DoGE improves the generalization of the base model to any target data mixture. On the SlimPajama dataset, our base model gets better perplexity and few-shot reasoning accuracies across $6$ tasks compared to baseline methods. Moreover, aiming to generalize to out-of-domain target tasks, which is unseen in the pretraining corpus (OOD domain), DoGE can effectively identify inter-domain dependencies, and consistently achieves better test perplexity on the target domain.
format Preprint
id arxiv_https___arxiv_org_abs_2310_15393
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DoGE: Domain Reweighting with Generalization Estimation
Fan, Simin
Pagliardini, Matteo
Jaggi, Martin
Machine Learning
Artificial Intelligence
Computation and Language
The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs). Despite its importance, recent LLMs still rely on heuristics and trial and error to increase or reduce the influence of data-domains. We propose DOmain reweighting with Generalization Estimation (DoGE), which optimizes the probability of sampling from each domain (domain weights) in a principled way. Our approach is a two-stage process consisting of (i) training a proxy model to obtain domain weights using a bi-level optimization algorithm; (ii) training a larger base model by sampling training domains according to the learned domain weights. In our experiments, we extensively show how DoGE improves the generalization of the base model to any target data mixture. On the SlimPajama dataset, our base model gets better perplexity and few-shot reasoning accuracies across $6$ tasks compared to baseline methods. Moreover, aiming to generalize to out-of-domain target tasks, which is unseen in the pretraining corpus (OOD domain), DoGE can effectively identify inter-domain dependencies, and consistently achieves better test perplexity on the target domain.
title DoGE: Domain Reweighting with Generalization Estimation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2310.15393