Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Grangier, David, Fan, Simin, Seto, Skyler, Ablin, Pierre
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913729458733056
author Grangier, David
Fan, Simin
Seto, Skyler
Ablin, Pierre
author_facet Grangier, David
Fan, Simin
Seto, Skyler
Ablin, Pierre
contents Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain these models is only available in limited amount for most tasks. In this work, we build specialist models from large generalist training sets instead. We propose a novel method, ClusteRed Importance SamPling (CRISP). CRISP clusters the generalist dataset and samples from these clusters based on their frequencies in the smaller specialist dataset. It is scalable, suitable for both pretraining and continued pretraining, and works well in multi-task settings. CRISP performs favorably compared to other methods that adjust the training distribution of the generalist data with guidance from the limited domain-specific data. Our findings demonstrate improvements across different domains in terms of language modeling perplexity and accuracy on multiple-choice question tasks. We also present ablation studies that examine the impact of dataset sizes, clustering configurations, and model sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03735
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
Grangier, David
Fan, Simin
Seto, Skyler
Ablin, Pierre
Computation and Language
Machine Learning
Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain these models is only available in limited amount for most tasks. In this work, we build specialist models from large generalist training sets instead. We propose a novel method, ClusteRed Importance SamPling (CRISP). CRISP clusters the generalist dataset and samples from these clusters based on their frequencies in the smaller specialist dataset. It is scalable, suitable for both pretraining and continued pretraining, and works well in multi-task settings. CRISP performs favorably compared to other methods that adjust the training distribution of the generalist data with guidance from the limited domain-specific data. Our findings demonstrate improvements across different domains in terms of language modeling perplexity and accuracy on multiple-choice question tasks. We also present ablation studies that examine the impact of dataset sizes, clustering configurations, and model sizes.
title Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.03735