From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Chejian, Ping, Wei, Xu, Peng, Liu, Zihan, Wang, Boxin, Shoeybi, Mohammad, Li, Bo, Catanzaro, Bryan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913783639703552
author Xu, Chejian
Ping, Wei
Xu, Peng
Liu, Zihan
Wang, Boxin
Shoeybi, Mohammad
Li, Bo
Catanzaro, Bryan
author_facet Xu, Chejian
Ping, Wei
Xu, Peng
Liu, Zihan
Wang, Boxin
Shoeybi, Mohammad
Li, Bo
Catanzaro, Bryan
contents Long-context capabilities are essential for a wide range of applications, including document and video understanding, in-context learning, and inference-time scaling, all of which require models to process and reason over long sequences of text and multimodal data. In this work, we introduce a efficient training recipe for building ultra-long context LLMs from aligned instruct model, pushing the boundaries of context lengths from 128K to 1M, 2M, and 4M tokens. Our approach leverages efficient continued pretraining strategies to extend the context window and employs effective instruction tuning to maintain the instruction-following and reasoning abilities. Our UltraLong-8B, built on Llama3.1-Instruct with our recipe, achieves state-of-the-art performance across a diverse set of long-context benchmarks. Importantly, models trained with our approach maintain competitive performance on standard benchmarks, demonstrating balanced improvements for both long and short context tasks. We further provide an in-depth analysis of key design choices, highlighting the impacts of scaling strategies and data composition. Our findings establish a robust framework for efficiently scaling context lengths while preserving general model capabilities. We release all model weights at: https://ultralong.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06214
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
Xu, Chejian
Ping, Wei
Xu, Peng
Liu, Zihan
Wang, Boxin
Shoeybi, Mohammad
Li, Bo
Catanzaro, Bryan
Computation and Language
Artificial Intelligence
Machine Learning
Long-context capabilities are essential for a wide range of applications, including document and video understanding, in-context learning, and inference-time scaling, all of which require models to process and reason over long sequences of text and multimodal data. In this work, we introduce a efficient training recipe for building ultra-long context LLMs from aligned instruct model, pushing the boundaries of context lengths from 128K to 1M, 2M, and 4M tokens. Our approach leverages efficient continued pretraining strategies to extend the context window and employs effective instruction tuning to maintain the instruction-following and reasoning abilities. Our UltraLong-8B, built on Llama3.1-Instruct with our recipe, achieves state-of-the-art performance across a diverse set of long-context benchmarks. Importantly, models trained with our approach maintain competitive performance on standard benchmarks, demonstrating balanced improvements for both long and short context tasks. We further provide an in-depth analysis of key design choices, highlighting the impacts of scaling strategies and data composition. Our findings establish a robust framework for efficiently scaling context lengths while preserving general model capabilities. We release all model weights at: https://ultralong.github.io/.
title From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.06214