SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913058166669312 |
|---|---|
| author | Xu, Hongtao Tan, Jianchao Hu, Yuxuan Lu, Pengju Wang, Hongyu Sun, Pingwei Sun, Yerui Xie, Yuchen Cai, Xunliang Li, Mingzhen Jia, Weile |
| author_facet | Xu, Hongtao Tan, Jianchao Hu, Yuxuan Lu, Pengju Wang, Hongyu Sun, Pingwei Sun, Yerui Xie, Yuchen Cai, Xunliang Li, Mingzhen Jia, Weile |
| contents | While sparse attention mitigates the computational bottleneck of long-context LLM training, its distributed training process exhibits extreme heterogeneity in both \textit{1)} sequence length and \textit{2)} sparsity sensitivity, leading to a severe imbalance problem and sub-optimal model accuracy. Existing algorithms and training frameworks typically focus on single issue, failing to systematically co-optimize these two problems. Therefore, we propose SparseBalance, a novel algorithm-system co-design framework, which exploits the sparsity and sequence heterogeneity to optimize model accuracy and system efficiency jointly. First, we propose workload-aware dynamic sparsity tuning, which employs a bidirectional sparsity adjustment to eliminate stragglers and exploit inherent bubbles for free accuracy. Second, we propose a sparsity-aware batching strategy to achieve coarse-grained balance, which complements dynamic sparsity tuning. Experimental results demonstrate that SparseBalance achieves up to a 1.33$\times$ end-to-end speedup while still improving the long-context capability by 0.46\% on the LongBench benchmark. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_13847 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention Xu, Hongtao Tan, Jianchao Hu, Yuxuan Lu, Pengju Wang, Hongyu Sun, Pingwei Sun, Yerui Xie, Yuchen Cai, Xunliang Li, Mingzhen Jia, Weile Machine Learning Artificial Intelligence While sparse attention mitigates the computational bottleneck of long-context LLM training, its distributed training process exhibits extreme heterogeneity in both \textit{1)} sequence length and \textit{2)} sparsity sensitivity, leading to a severe imbalance problem and sub-optimal model accuracy. Existing algorithms and training frameworks typically focus on single issue, failing to systematically co-optimize these two problems. Therefore, we propose SparseBalance, a novel algorithm-system co-design framework, which exploits the sparsity and sequence heterogeneity to optimize model accuracy and system efficiency jointly. First, we propose workload-aware dynamic sparsity tuning, which employs a bidirectional sparsity adjustment to eliminate stragglers and exploit inherent bubbles for free accuracy. Second, we propose a sparsity-aware batching strategy to achieve coarse-grained balance, which complements dynamic sparsity tuning. Experimental results demonstrate that SparseBalance achieves up to a 1.33$\times$ end-to-end speedup while still improving the long-context capability by 0.46\% on the LongBench benchmark. |
| title | SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2604.13847 |