LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Minghao, Waheed, Abdul, Zhang, Chiyu, Abdul-Mageed, Muhammad, Aji, Alham Fikri
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911765720203264
author Wu, Minghao
Waheed, Abdul
Zhang, Chiyu
Abdul-Mageed, Muhammad
Aji, Alham Fikri
author_facet Wu, Minghao
Waheed, Abdul
Zhang, Chiyu
Abdul-Mageed, Muhammad
Aji, Alham Fikri
contents Large language models (LLMs) with instruction fine-tuning demonstrate superior generative capabilities. However, these models are resource-intensive. To alleviate this issue, we explore distilling knowledge from instruction-tuned LLMs into much smaller ones. To this end, we carefully develop a large set of 2.58M instructions based on both existing and newly-generated instructions. In addition to being sizable, we design our instructions to cover a broad set of topics to ensure diversity. Extensive analysis of our instruction dataset confirms its diversity, and we generate responses for these instructions using gpt-3.5-turbo. Leveraging these instructions, we fine-tune a diverse herd of models, collectively referred to as LaMini-LM, which includes models from both the encoder-decoder and decoder-only families, with varying sizes. We evaluate the performance of our models using automatic metrics on 15 different natural language processing (NLP) benchmarks, as well as through human assessment. The results demonstrate that our proposed LaMini-LM models are comparable to competitive baselines, while being much smaller in size.
format Preprint
id arxiv_https___arxiv_org_abs_2304_14402
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
Wu, Minghao
Waheed, Abdul
Zhang, Chiyu
Abdul-Mageed, Muhammad
Aji, Alham Fikri
Computation and Language
Large language models (LLMs) with instruction fine-tuning demonstrate superior generative capabilities. However, these models are resource-intensive. To alleviate this issue, we explore distilling knowledge from instruction-tuned LLMs into much smaller ones. To this end, we carefully develop a large set of 2.58M instructions based on both existing and newly-generated instructions. In addition to being sizable, we design our instructions to cover a broad set of topics to ensure diversity. Extensive analysis of our instruction dataset confirms its diversity, and we generate responses for these instructions using gpt-3.5-turbo. Leveraging these instructions, we fine-tune a diverse herd of models, collectively referred to as LaMini-LM, which includes models from both the encoder-decoder and decoder-only families, with varying sizes. We evaluate the performance of our models using automatic metrics on 15 different natural language processing (NLP) benchmarks, as well as through human assessment. The results demonstrate that our proposed LaMini-LM models are comparable to competitive baselines, while being much smaller in size.
title LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
topic Computation and Language
url https://arxiv.org/abs/2304.14402