MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Wenxuan, Zhang, Chengruidong, Jiang, Huiqiang, Li, Yucheng, Yang, Yuqing, Qiu, Lili
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914579809828864
author Li, Wenxuan
Zhang, Chengruidong
Jiang, Huiqiang
Li, Yucheng
Yang, Yuqing
Qiu, Lili
author_facet Li, Wenxuan
Zhang, Chengruidong
Jiang, Huiqiang
Li, Yucheng
Yang, Yuqing
Qiu, Lili
contents The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their capacity for complex reasoning and broaden their applicability across diverse scenarios. Dynamic sparse attention is a promising approach for reducing the computational cost of long-context. However, efficiently training LLMs with dynamic sparse attention on ultra-long contexts-especially in distributed settings-remains a significant challenge, due in large part to worker- and step-level imbalance. This paper introduces MTraining, a novel distributed methodology leveraging dynamic sparse attention to enable efficient training for LLMs with ultra-long contexts. Specifically, MTraining integrates three key components: a dynamic sparse training pattern, balanced sparse ring attention, and hierarchical sparse ring attention. These components are designed to synergistically address the computational imbalance and communication overheads inherent in dynamic sparse attention mechanisms during the training of models with extensive context lengths. We demonstrate the efficacy of MTraining by training Qwen2.5-3B, successfully expanding its context window from 32K to 512K tokens on a cluster of 32 A100 GPUs. Our evaluations on a comprehensive suite of downstream tasks, including RULER, PG-19, InfiniteBench, and Needle In A Haystack, reveal that MTraining achieves up to a 6x higher training throughput while preserving model accuracy. Our code is available at https://github.com/microsoft/MInference/tree/main/MTraining.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18830
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
Li, Wenxuan
Zhang, Chengruidong
Jiang, Huiqiang
Li, Yucheng
Yang, Yuqing
Qiu, Lili
Computation and Language
Distributed, Parallel, and Cluster Computing
Machine Learning
The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their capacity for complex reasoning and broaden their applicability across diverse scenarios. Dynamic sparse attention is a promising approach for reducing the computational cost of long-context. However, efficiently training LLMs with dynamic sparse attention on ultra-long contexts-especially in distributed settings-remains a significant challenge, due in large part to worker- and step-level imbalance. This paper introduces MTraining, a novel distributed methodology leveraging dynamic sparse attention to enable efficient training for LLMs with ultra-long contexts. Specifically, MTraining integrates three key components: a dynamic sparse training pattern, balanced sparse ring attention, and hierarchical sparse ring attention. These components are designed to synergistically address the computational imbalance and communication overheads inherent in dynamic sparse attention mechanisms during the training of models with extensive context lengths. We demonstrate the efficacy of MTraining by training Qwen2.5-3B, successfully expanding its context window from 32K to 512K tokens on a cluster of 32 A100 GPUs. Our evaluations on a comprehensive suite of downstream tasks, including RULER, PG-19, InfiniteBench, and Needle In A Haystack, reveal that MTraining achieves up to a 6x higher training throughput while preserving model accuracy. Our code is available at https://github.com/microsoft/MInference/tree/main/MTraining.
title MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
topic Computation and Language
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2510.18830