HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Ran, Jiang, Youhe, Nie, Xiaonan, Fu, Fangcheng, Cui, Bin, Yuan, Binhang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
Cascadia: An Efficient Cascade Serving System for Large Language Models
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
Improving Automatic Parallel Training via Balanced Memory Workload Optimization
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
von: Lin, Sheng, et al.
Veröffentlicht: (2025)
von: Lin, Sheng, et al.
Veröffentlicht: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
von: Tong, Chris, et al.
Veröffentlicht: (2025)
von: Tong, Chris, et al.
Veröffentlicht: (2025)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
von: Ge, Hao, et al.
Veröffentlicht: (2025)
von: Ge, Hao, et al.
Veröffentlicht: (2025)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
von: Zhao, Pinxue, et al.
Veröffentlicht: (2024)
von: Zhao, Pinxue, et al.
Veröffentlicht: (2024)
FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
von: Wang, Yujie, et al.
Veröffentlicht: (2024)
von: Wang, Yujie, et al.
Veröffentlicht: (2024)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
von: Tang, Ding, et al.
Veröffentlicht: (2025)
von: Tang, Ding, et al.
Veröffentlicht: (2025)
Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
von: Wang, Yujie, et al.
Veröffentlicht: (2024)
von: Wang, Yujie, et al.
Veröffentlicht: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
von: Fernandez, Jared, et al.
Veröffentlicht: (2024)
von: Fernandez, Jared, et al.
Veröffentlicht: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
von: Li, Pengbo, et al.
Veröffentlicht: (2026)
von: Li, Pengbo, et al.
Veröffentlicht: (2026)
Revisiting the Time Cost Model of AllReduce
von: Xiong, Dian, et al.
Veröffentlicht: (2024)
von: Xiong, Dian, et al.
Veröffentlicht: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
von: Hu, Rizhen, et al.
Veröffentlicht: (2025)
von: Hu, Rizhen, et al.
Veröffentlicht: (2025)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
von: Xu, Si, et al.
Veröffentlicht: (2024)
von: Xu, Si, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026) -
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025) -
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023) -
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025) -
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)