LongAlign: A Recipe for Long Context Alignment of Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Bai, Yushi, Lv, Xin, Zhang, Jiajie, He, Yuze, Qi, Ji, Hou, Lei, Tang, Jie, Dong, Yuxiao, Li, Juanzi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917579728093184
author Bai, Yushi
Lv, Xin
Zhang, Jiajie
He, Yuze
Qi, Ji
Hou, Lei
Tang, Jie
Dong, Yuxiao
Li, Juanzi
author_facet Bai, Yushi
Lv, Xin
Zhang, Jiajie
He, Yuze
Qi, Ji
Hou, Lei
Tang, Jie
Dong, Yuxiao
Li, Juanzi
contents Extending large language models to effectively handle long contexts requires instruction fine-tuning on input sequences of similar length. To address this, we present LongAlign -- a recipe of the instruction data, training, and evaluation for long context alignment. First, we construct a long instruction-following dataset using Self-Instruct. To ensure the data diversity, it covers a broad range of tasks from various long context sources. Second, we adopt the packing and sorted batching strategies to speed up supervised fine-tuning on data with varied length distributions. Additionally, we develop a loss weighting method to balance the contribution to the loss across different sequences during packing training. Third, we introduce the LongBench-Chat benchmark for evaluating instruction-following capabilities on queries of 10k-100k in length. Experiments show that LongAlign outperforms existing recipes for LLMs in long context tasks by up to 30\%, while also maintaining their proficiency in handling short, generic tasks. The code, data, and long-aligned models are open-sourced at https://github.com/THUDM/LongAlign.
format Preprint
id arxiv_https___arxiv_org_abs_2401_18058
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LongAlign: A Recipe for Long Context Alignment of Large Language Models
Bai, Yushi
Lv, Xin
Zhang, Jiajie
He, Yuze
Qi, Ji
Hou, Lei
Tang, Jie
Dong, Yuxiao
Li, Juanzi
Computation and Language
Machine Learning
Extending large language models to effectively handle long contexts requires instruction fine-tuning on input sequences of similar length. To address this, we present LongAlign -- a recipe of the instruction data, training, and evaluation for long context alignment. First, we construct a long instruction-following dataset using Self-Instruct. To ensure the data diversity, it covers a broad range of tasks from various long context sources. Second, we adopt the packing and sorted batching strategies to speed up supervised fine-tuning on data with varied length distributions. Additionally, we develop a loss weighting method to balance the contribution to the loss across different sequences during packing training. Third, we introduce the LongBench-Chat benchmark for evaluating instruction-following capabilities on queries of 10k-100k in length. Experiments show that LongAlign outperforms existing recipes for LLMs in long context tasks by up to 30\%, while also maintaining their proficiency in handling short, generic tasks. The code, data, and long-aligned models are open-sourced at https://github.com/THUDM/LongAlign.
title LongAlign: A Recipe for Long Context Alignment of Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2401.18058