Efficient Long-context Language Model Training by Core Attention Disaggregation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhuang, Yonghao, Chen, Junda, Pang, Bo, Gu, Yi, Zhu, Yibo, Jiang, Yimin, Stoica, Ion, Xing, Eric, Zhang, Hao |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
par: Li, Dacheng, et autres
Publié: (2023)
par: Li, Dacheng, et autres
Publié: (2023)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
par: Zhang, Zili, et autres
Publié: (2024)
par: Zhang, Zili, et autres
Publié: (2024)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
par: Zhong, Yinmin, et autres
Publié: (2024)
par: Zhong, Yinmin, et autres
Publié: (2024)
On Optimizing the Communication of Model Parallelism
par: Zhuang, Yonghao, et autres
Publié: (2022)
par: Zhuang, Yonghao, et autres
Publié: (2022)
The Dawn of Disaggregation and the Coherence Conundrum: A Call for Federated Coherence
par: Hong, Jaewan, et autres
Publié: (2025)
par: Hong, Jaewan, et autres
Publié: (2025)
Accelerating Distributed MoE Training and Inference with Lina
par: Li, Jiamin, et autres
Publié: (2022)
par: Li, Jiamin, et autres
Publié: (2022)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
par: Tan, Xin, et autres
Publié: (2026)
par: Tan, Xin, et autres
Publié: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
par: Zhang, Zhexiang, et autres
Publié: (2025)
par: Zhang, Zhexiang, et autres
Publié: (2025)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
par: Zhao, Yilong, et autres
Publié: (2026)
par: Zhao, Yilong, et autres
Publié: (2026)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
par: Kumar, Madabattula Rajesh, et autres
Publié: (2025)
par: Kumar, Madabattula Rajesh, et autres
Publié: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
par: Xue, Zhenliang, et autres
Publié: (2025)
par: Xue, Zhenliang, et autres
Publié: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
par: Liu, Guowei, et autres
Publié: (2026)
par: Liu, Guowei, et autres
Publié: (2026)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
par: Chen, Shaoyuan, et autres
Publié: (2024)
par: Chen, Shaoyuan, et autres
Publié: (2024)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
par: Hu, Zhisheng, et autres
Publié: (2025)
par: Hu, Zhisheng, et autres
Publié: (2025)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
par: Zhong, Yinmin, et autres
Publié: (2025)
par: Zhong, Yinmin, et autres
Publié: (2025)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
par: Wu, Tianyuan, et autres
Publié: (2025)
par: Wu, Tianyuan, et autres
Publié: (2025)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
par: Xia, Tian, et autres
Publié: (2025)
par: Xia, Tian, et autres
Publié: (2025)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
par: Gu, Diandian, et autres
Publié: (2024)
par: Gu, Diandian, et autres
Publié: (2024)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
par: Jiang, Youhe, et autres
Publié: (2025)
par: Jiang, Youhe, et autres
Publié: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
Pie: Pooling CPU Memory for LLM Inference
par: Xu, Yi, et autres
Publié: (2024)
par: Xu, Yi, et autres
Publié: (2024)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
par: Wu, Yongji, et autres
Publié: (2025)
par: Wu, Yongji, et autres
Publié: (2025)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
par: Xing, Jiarong, et autres
Publié: (2025)
par: Xing, Jiarong, et autres
Publié: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
par: Liu, Yunzhao, et autres
Publié: (2025)
par: Liu, Yunzhao, et autres
Publié: (2025)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
par: Wu, Yongji, et autres
Publié: (2025)
par: Wu, Yongji, et autres
Publié: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
par: Chen, Xing, et autres
Publié: (2025)
par: Chen, Xing, et autres
Publié: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
par: Basit, Omar, et autres
Publié: (2026)
par: Basit, Omar, et autres
Publié: (2026)
Software Resource Disaggregation for HPC with Serverless Computing
par: Copik, Marcin, et autres
Publié: (2024)
par: Copik, Marcin, et autres
Publié: (2024)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
par: Yang, Shuo, et autres
Publié: (2026)
par: Yang, Shuo, et autres
Publié: (2026)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
DEX: Scalable Range Indexing on Disaggregated Memory [Extended Version]
par: Lu, Baotong, et autres
Publié: (2024)
par: Lu, Baotong, et autres
Publié: (2024)
DRackSim: Simulator for Rack-scale Memory Disaggregation
par: Puri, Amit, et autres
Publié: (2023)
par: Puri, Amit, et autres
Publié: (2023)
Towards Disaggregation-Native Data Streaming between Devices
par: Asmussen, Nils, et autres
Publié: (2024)
par: Asmussen, Nils, et autres
Publié: (2024)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
par: Murat, Antoine, et autres
Publié: (2024)
par: Murat, Antoine, et autres
Publié: (2024)
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
par: Chen, Qiaoling, et autres
Publié: (2024)
par: Chen, Qiaoling, et autres
Publié: (2024)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
par: Ren, Feng, et autres
Publié: (2026)
par: Ren, Feng, et autres
Publié: (2026)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Documents similaires
-
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
par: Li, Dacheng, et autres
Publié: (2023) -
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
par: Zhang, Zili, et autres
Publié: (2024) -
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
par: Zhong, Yinmin, et autres
Publié: (2024) -
On Optimizing the Communication of Model Parallelism
par: Zhuang, Yonghao, et autres
Publié: (2022) -
The Dawn of Disaggregation and the Coherence Conundrum: A Call for Federated Coherence
par: Hong, Jaewan, et autres
Publié: (2025)