SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Hongtao, Tan, Jianchao, Hu, Yuxuan, Lu, Pengju, Wang, Hongyu, Sun, Pingwei, Sun, Yerui, Xie, Yuchen, Cai, Xunliang, Li, Mingzhen, Jia, Weile |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
by: Hu, Yuxuan, et al.
Published: (2026)
by: Hu, Yuxuan, et al.
Published: (2026)
FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
by: Sun, Pingwei, et al.
Published: (2026)
by: Sun, Pingwei, et al.
Published: (2026)
Accelerate Speculative Decoding with Sparse Computation in Verification
by: Wang, Jikai, et al.
Published: (2025)
by: Wang, Jikai, et al.
Published: (2025)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
by: Li, Jiacheng, et al.
Published: (2026)
by: Li, Jiacheng, et al.
Published: (2026)
WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
by: Li, Jiacheng, et al.
Published: (2025)
by: Li, Jiacheng, et al.
Published: (2025)
AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
by: Li, Jiacheng, et al.
Published: (2025)
by: Li, Jiacheng, et al.
Published: (2025)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
by: Wang, Hongyu, et al.
Published: (2026)
by: Wang, Hongyu, et al.
Published: (2026)
Large-scale Neural Network Quantum States for ab initio Quantum Chemistry Simulations on Fugaku
by: Xu, Hongtao, et al.
Published: (2025)
by: Xu, Hongtao, et al.
Published: (2025)
Load Balancing Using Sparse Communication
by: Mendelson, Gal, et al.
Published: (2022)
by: Mendelson, Gal, et al.
Published: (2022)
Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data Scheduling
by: Xu, Hongtao, et al.
Published: (2025)
by: Xu, Hongtao, et al.
Published: (2025)
Lag-Relative Sparse Attention In Long Context Training
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
by: Liu, Anmin, et al.
Published: (2026)
by: Liu, Anmin, et al.
Published: (2026)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025)
by: Vasylenko, Pavlo, et al.
Published: (2025)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
by: Tang, Hongyin, et al.
Published: (2024)
by: Tang, Hongyin, et al.
Published: (2024)
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
by: Huo, Feiye, et al.
Published: (2025)
by: Huo, Feiye, et al.
Published: (2025)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Partitioning Unstructured Sparse Tensor Algebra for Load-Balanced Parallel Execution
by: Chougule, Atharva, et al.
Published: (2026)
by: Chougule, Atharva, et al.
Published: (2026)
Sparse Mean Field Load Balancing in Large Localized Queueing Systems
by: Tahir, Anam, et al.
Published: (2023)
by: Tahir, Anam, et al.
Published: (2023)
Efficient Context Scaling with LongCat ZigZag Attention
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026)
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026)
Embedded Federated Feature Selection with Dynamic Sparse Training: Balancing Accuracy-Cost Tradeoffs
by: Mahanipour, Afsaneh, et al.
Published: (2025)
by: Mahanipour, Afsaneh, et al.
Published: (2025)
Fine-tuning vs Prompting, Can Language Models Understand Human Values?
by: Sun, Pingwei
Published: (2024)
by: Sun, Pingwei
Published: (2024)
Efficient Long Context Fine-tuning with Chunk Flow
by: Yuan, Xiulong, et al.
Published: (2025)
by: Yuan, Xiulong, et al.
Published: (2025)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
by: Liu, Siran, et al.
Published: (2026)
by: Liu, Siran, et al.
Published: (2026)
db-SP: Accelerating Sparse Attention for Visual Generative Models with Dual-Balanced Sequence Parallelism
by: Chen, Siqi, et al.
Published: (2025)
by: Chen, Siqi, et al.
Published: (2025)
Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs
by: Zhang, Yuxin, et al.
Published: (2023)
by: Zhang, Yuxin, et al.
Published: (2023)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs
by: Ni, Wentao, et al.
Published: (2026)
by: Ni, Wentao, et al.
Published: (2026)
Gated Sparse Attention: Combining Computational Efficiency with Training Stability for Long-Context Language Models
by: Shen, Alfred, et al.
Published: (2026)
by: Shen, Alfred, et al.
Published: (2026)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
Expert Threshold Routing for Autoregressive Language Modeling with Dynamic Computation Allocation and Load Balancing
by: Sun, Hanchi, et al.
Published: (2026)
by: Sun, Hanchi, et al.
Published: (2026)
Load Balancing for AI Training Workloads
by: McClure, Sarah, et al.
Published: (2025)
by: McClure, Sarah, et al.
Published: (2025)
A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs
by: Liu, Zijie, et al.
Published: (2026)
by: Liu, Zijie, et al.
Published: (2026)
A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention
by: Xiu, Di, et al.
Published: (2025)
by: Xiu, Di, et al.
Published: (2025)
LVSA: Training-Free Sparse Attention for Long Video Diffusion
by: Glorian, Gael, et al.
Published: (2026)
by: Glorian, Gael, et al.
Published: (2026)
Similar Items
-
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
by: Hu, Yuxuan, et al.
Published: (2025) -
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
by: Hu, Yuxuan, et al.
Published: (2026) -
FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
by: Sun, Pingwei, et al.
Published: (2026) -
Accelerate Speculative Decoding with Sparse Computation in Verification
by: Wang, Jikai, et al.
Published: (2025) -
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
by: Li, Jiacheng, et al.
Published: (2026)