DEPT: Decoupled Embeddings for Pre-training Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Iacob, Alex, Sani, Lorenzo, Kurmanji, Meghdad, Shen, William F., Qiu, Xinchi, Cai, Dongqi, Gao, Yan, Lane, Nicholas D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909567861915648
author Iacob, Alex
Sani, Lorenzo
Kurmanji, Meghdad
Shen, William F.
Qiu, Xinchi
Cai, Dongqi
Gao, Yan
Lane, Nicholas D.
author_facet Iacob, Alex
Sani, Lorenzo
Kurmanji, Meghdad
Shen, William F.
Qiu, Xinchi
Cai, Dongqi
Gao, Yan
Lane, Nicholas D.
contents Language Model pre-training uses broad data mixtures to enhance performance across domains and languages. However, training on such heterogeneous text corpora requires extensive and expensive efforts. Since these data sources vary significantly in lexical, syntactic, and semantic aspects, they cause negative interference or the ``curse of multilinguality''. To address these challenges we propose a communication-efficient pre-training framework, DEPT. Our method decouples embeddings from the transformer body while simultaneously training the latter on multiple data sources without requiring a shared vocabulary. DEPT can: (1) train robustly and effectively under significant data heterogeneity, (2) minimize token embedding parameters to only what the data source vocabulary requires, while cutting communication costs in direct proportion to both the communication frequency and the reduction in parameters, (3) enhance transformer body plasticity and generalization, improving both average perplexity (up to 20%) and downstream task performance, and (4) enable training with custom optimized vocabularies per data source. We demonstrate DEPT's potential via the first vocabulary-agnostic federated pre-training of billion-scale models, reducing communication costs by orders of magnitude and embedding memory by 4-5x.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05021
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DEPT: Decoupled Embeddings for Pre-training Language Models
Iacob, Alex
Sani, Lorenzo
Kurmanji, Meghdad
Shen, William F.
Qiu, Xinchi
Cai, Dongqi
Gao, Yan
Lane, Nicholas D.
Machine Learning
Computation and Language
Language Model pre-training uses broad data mixtures to enhance performance across domains and languages. However, training on such heterogeneous text corpora requires extensive and expensive efforts. Since these data sources vary significantly in lexical, syntactic, and semantic aspects, they cause negative interference or the ``curse of multilinguality''. To address these challenges we propose a communication-efficient pre-training framework, DEPT. Our method decouples embeddings from the transformer body while simultaneously training the latter on multiple data sources without requiring a shared vocabulary. DEPT can: (1) train robustly and effectively under significant data heterogeneity, (2) minimize token embedding parameters to only what the data source vocabulary requires, while cutting communication costs in direct proportion to both the communication frequency and the reduction in parameters, (3) enhance transformer body plasticity and generalization, improving both average perplexity (up to 20%) and downstream task performance, and (4) enable training with custom optimized vocabularies per data source. We demonstrate DEPT's potential via the first vocabulary-agnostic federated pre-training of billion-scale models, reducing communication costs by orders of magnitude and embedding memory by 4-5x.
title DEPT: Decoupled Embeddings for Pre-training Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2410.05021