DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Dacheng, Shao, Rulin, Xie, Anze, Xing, Eric P., Ma, Xuezhe, Stoica, Ion, Gonzalez, Joseph E., Zhang, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Long-context Language Model Training by Core Attention Disaggregation
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2025)
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2025)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
On Optimizing the Communication of Model Parallelism
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2022)
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2022)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
von: Xia, Tian, et al.
Veröffentlicht: (2025)
von: Xia, Tian, et al.
Veröffentlicht: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
Implementing OpenMP for Zig to enable its use in HPC context
von: Kacs, David, et al.
Veröffentlicht: (2024)
von: Kacs, David, et al.
Veröffentlicht: (2024)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
von: Li, Cong, et al.
Veröffentlicht: (2025)
von: Li, Cong, et al.
Veröffentlicht: (2025)
Locality-aware Fair Scheduling in LLM Serving
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
von: Sun, Ao, et al.
Veröffentlicht: (2025)
von: Sun, Ao, et al.
Veröffentlicht: (2025)
SkyStore: Cost-Optimized Object Storage Across Regions and Clouds
von: Liu, Shu, et al.
Veröffentlicht: (2025)
von: Liu, Shu, et al.
Veröffentlicht: (2025)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
LOCO: Rethinking Objects for Network Memory
von: Hodgkins, George, et al.
Veröffentlicht: (2025)
von: Hodgkins, George, et al.
Veröffentlicht: (2025)
D-CAST: Distributed Consensus Switch in Wireless Trustworthy Autonomous System
von: Yu, Dachao, et al.
Veröffentlicht: (2024)
von: Yu, Dachao, et al.
Veröffentlicht: (2024)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
von: Du, Kuntai, et al.
Veröffentlicht: (2025)
von: Du, Kuntai, et al.
Veröffentlicht: (2025)
Federated Neural Radiance Field for Distributed Intelligence
von: Zhang, Yintian, et al.
Veröffentlicht: (2024)
von: Zhang, Yintian, et al.
Veröffentlicht: (2024)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
von: Liu, Xueshen, et al.
Veröffentlicht: (2026)
von: Liu, Xueshen, et al.
Veröffentlicht: (2026)
The Dawn of Disaggregation and the Coherence Conundrum: A Call for Federated Coherence
von: Hong, Jaewan, et al.
Veröffentlicht: (2025)
von: Hong, Jaewan, et al.
Veröffentlicht: (2025)
Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling
von: Bader, Jonathan, et al.
Veröffentlicht: (2026)
von: Bader, Jonathan, et al.
Veröffentlicht: (2026)
Solutions for Distributed Memory Access Mechanism on HPC Clusters
von: Meizner, Jan, et al.
Veröffentlicht: (2025)
von: Meizner, Jan, et al.
Veröffentlicht: (2025)
Efficient Distributed MLLM Training with Cornstarch
von: Jang, Insu, et al.
Veröffentlicht: (2025)
von: Jang, Insu, et al.
Veröffentlicht: (2025)
DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
von: Ma, Haoran, et al.
Veröffentlicht: (2024)
von: Ma, Haoran, et al.
Veröffentlicht: (2024)
Distributed Inference Performance Optimization for LLMs on CPUs
von: He, Pujiang, et al.
Veröffentlicht: (2024)
von: He, Pujiang, et al.
Veröffentlicht: (2024)
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
von: Wang, Haiquan, et al.
Veröffentlicht: (2024)
von: Wang, Haiquan, et al.
Veröffentlicht: (2024)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient Long-context Language Model Training by Core Attention Disaggregation
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2025) -
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
von: Zhao, Yilong, et al.
Veröffentlicht: (2026) -
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024) -
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
von: Zhang, Chen, et al.
Veröffentlicht: (2025) -
On Optimizing the Communication of Model Parallelism
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2022)