LEGO: Language Model Building Blocks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhansali, Shrenik, Jin, Alwin, Lizzo, Tyler, Heck, Larry
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913561620512768
author Bhansali, Shrenik
Jin, Alwin
Lizzo, Tyler
Heck, Larry
author_facet Bhansali, Shrenik
Jin, Alwin
Lizzo, Tyler
Heck, Larry
contents Large language models (LLMs) are essential in natural language processing (NLP) but are costly in data collection, pre-training, fine-tuning, and inference. Task-specific small language models (SLMs) offer a cheaper alternative but lack robustness and generalization. This paper proposes LEGO, a novel technique to extract SLMs from an LLM and recombine them. Using state-of-the-art LLM pruning strategies, we can create task- and user-specific SLM building blocks that are efficient for fine-tuning and inference while also preserving user data privacy. LEGO utilizes Federated Learning and a novel aggregation scheme for the LLM reconstruction, maintaining robustness without high costs and preserving user data privacy. We experimentally demonstrate the versatility of LEGO, showing its ability to enable model heterogeneity and mitigate the effects of data heterogeneity while maintaining LLM robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18287
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LEGO: Language Model Building Blocks
Bhansali, Shrenik
Jin, Alwin
Lizzo, Tyler
Heck, Larry
Computation and Language
Machine Learning
Large language models (LLMs) are essential in natural language processing (NLP) but are costly in data collection, pre-training, fine-tuning, and inference. Task-specific small language models (SLMs) offer a cheaper alternative but lack robustness and generalization. This paper proposes LEGO, a novel technique to extract SLMs from an LLM and recombine them. Using state-of-the-art LLM pruning strategies, we can create task- and user-specific SLM building blocks that are efficient for fine-tuning and inference while also preserving user data privacy. LEGO utilizes Federated Learning and a novel aggregation scheme for the LLM reconstruction, maintaining robustness without high costs and preserving user data privacy. We experimentally demonstrate the versatility of LEGO, showing its ability to enable model heterogeneity and mitigate the effects of data heterogeneity while maintaining LLM robustness.
title LEGO: Language Model Building Blocks
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.18287