Extending Llama-3's Context Ten-Fold Overnight

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Peitian, Shao, Ninglu, Liu, Zheng, Xiao, Shitao, Qian, Hongjin, Ye, Qiwei, Dou, Zhicheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914777718063104
author Zhang, Peitian
Shao, Ninglu
Liu, Zheng
Xiao, Shitao
Qian, Hongjin
Ye, Qiwei
Dou, Zhicheng
author_facet Zhang, Peitian
Shao, Ninglu
Liu, Zheng
Xiao, Shitao
Qian, Hongjin
Ye, Qiwei
Dou, Zhicheng
contents We extend the context length of Llama-3-8B-Instruct from 8K to 80K via QLoRA fine-tuning. The entire training cycle is super efficient, which takes 8 hours on one 8xA800 (80G) GPU machine. The resulted model exhibits superior performances across a broad range of evaluation tasks, such as NIHS, topic retrieval, and long-context language understanding; meanwhile, it also well preserves the original capability over short contexts. The dramatic context extension is mainly attributed to merely 3.5K synthetic training samples generated by GPT-4 , which indicates the LLMs' inherent (yet largely underestimated) potential to extend its original context length. In fact, the context length could be extended far beyond 80K with more computation resources. Therefore, the team will publicly release the entire resources (including data, model, data generation pipeline, training code) so as to facilitate the future research from the community: \url{https://github.com/FlagOpen/FlagEmbedding}.
format Preprint
id arxiv_https___arxiv_org_abs_2404_19553
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extending Llama-3's Context Ten-Fold Overnight
Zhang, Peitian
Shao, Ninglu
Liu, Zheng
Xiao, Shitao
Qian, Hongjin
Ye, Qiwei
Dou, Zhicheng
Computation and Language
We extend the context length of Llama-3-8B-Instruct from 8K to 80K via QLoRA fine-tuning. The entire training cycle is super efficient, which takes 8 hours on one 8xA800 (80G) GPU machine. The resulted model exhibits superior performances across a broad range of evaluation tasks, such as NIHS, topic retrieval, and long-context language understanding; meanwhile, it also well preserves the original capability over short contexts. The dramatic context extension is mainly attributed to merely 3.5K synthetic training samples generated by GPT-4 , which indicates the LLMs' inherent (yet largely underestimated) potential to extend its original context length. In fact, the context length could be extended far beyond 80K with more computation resources. Therefore, the team will publicly release the entire resources (including data, model, data generation pipeline, training code) so as to facilitate the future research from the community: \url{https://github.com/FlagOpen/FlagEmbedding}.
title Extending Llama-3's Context Ten-Fold Overnight
topic Computation and Language
url https://arxiv.org/abs/2404.19553