Gespeichert in:
| Hauptverfasser: | Wang, Zezhou, Li, Youjie, Lin, Zhiqi, Yang, Jiacheng, Xie, Cong, Feng, Guanyu, Zhong, Zheng, Huang, Ziyue, Zhu, Hongyu, Zhang, Zhi, Peng, Yanghua, Liu, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.22437 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
von: Li, Youjie, et al.
Veröffentlicht: (2025)
von: Li, Youjie, et al.
Veröffentlicht: (2025)
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
von: Zhang, Ruisi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruisi, et al.
Veröffentlicht: (2024)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025)
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025)
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
von: Ma, Qianli, et al.
Veröffentlicht: (2025)
von: Ma, Qianli, et al.
Veröffentlicht: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
Exploring Uncore Frequency Scaling for Heterogeneous Computing
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
von: Feng, Weiqi, et al.
Veröffentlicht: (2024)
von: Feng, Weiqi, et al.
Veröffentlicht: (2024)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
von: Hu, Zhisheng, et al.
Veröffentlicht: (2025)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
von: Wu, Xin, et al.
Veröffentlicht: (2026)
von: Wu, Xin, et al.
Veröffentlicht: (2026)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
von: Zhang, Zuoning, et al.
Veröffentlicht: (2024)
von: Zhang, Zuoning, et al.
Veröffentlicht: (2024)
Comparing Cross-Platform Performance via Node-to-Node Scaling Studies
von: Weiss, Kenneth, et al.
Veröffentlicht: (2025)
von: Weiss, Kenneth, et al.
Veröffentlicht: (2025)
Data Caching for Enterprise-Grade Petabyte-Scale OLAP
von: Tang, Chunxu, et al.
Veröffentlicht: (2024)
von: Tang, Chunxu, et al.
Veröffentlicht: (2024)
EdgeVision: Towards Collaborative Video Analytics on Distributed Edges for Performance Maximization
von: Gao, Guanyu, et al.
Veröffentlicht: (2022)
von: Gao, Guanyu, et al.
Veröffentlicht: (2022)
EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
madupite: A High-Performance Distributed Solver for Large-Scale Markov Decision Processes
von: Gargiani, Matilde, et al.
Veröffentlicht: (2025)
von: Gargiani, Matilde, et al.
Veröffentlicht: (2025)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
PolarStore: High-Performance Data Compression for Large-Scale Cloud-Native Databases
von: Hu, Qingda, et al.
Veröffentlicht: (2025)
von: Hu, Qingda, et al.
Veröffentlicht: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
M$^2$-MFP: A Multi-Scale and Multi-Level Memory Failure Prediction Framework for Reliable Cloud Infrastructure
von: Xie, Hongyi, et al.
Veröffentlicht: (2025)
von: Xie, Hongyi, et al.
Veröffentlicht: (2025)
Optimizing High-Throughput Distributed Data Pipelines for Reproducible Deep Learning at Scale
von: Mittal, Kashish, et al.
Veröffentlicht: (2026)
von: Mittal, Kashish, et al.
Veröffentlicht: (2026)
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
Case Study: Performance Analysis of a Virtualized XRootD Frontend in Large-Scale WAN Transfers
von: da Silva, J M, et al.
Veröffentlicht: (2026)
von: da Silva, J M, et al.
Veröffentlicht: (2026)
Deep Learning-Enabled Supercritical Flame Simulation at Detailed Chemistry and Real-Fluid Accuracy Towards Trillion-Cell Scale
von: Guo, Zhuoqiang, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoqiang, et al.
Veröffentlicht: (2025)
FAIR Ecosystems for Science at Scale
von: Wilkinson, Sean R., et al.
Veröffentlicht: (2025)
von: Wilkinson, Sean R., et al.
Veröffentlicht: (2025)
Scaling MPI Applications on Aurora
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
Steering a Fleet: Adaptation for Large-Scale, Workflow-Based Experiments
von: Pruyne, Jim, et al.
Veröffentlicht: (2024)
von: Pruyne, Jim, et al.
Veröffentlicht: (2024)
MPI-Q: A Message Communication Library for Large-Scale Classical-Quantum Heterogeneous Hybrid Distributed Computing
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
Scaling Real-Time Traffic Analytics on Edge-Cloud Fabrics for City-Scale Camera Networks
von: Sharma, Akash, et al.
Veröffentlicht: (2026)
von: Sharma, Akash, et al.
Veröffentlicht: (2026)
SDSL-Solver: Scalable Distributed Sparse Linear Solvers for Large-Scale Interior Point Methods
von: Yang, Shaofeng, et al.
Veröffentlicht: (2026)
von: Yang, Shaofeng, et al.
Veröffentlicht: (2026)
Barycentric Coded Distributed Computing with Flexible Recovery Threshold for Collaborative Mobile Edge Computing
von: Qiu, Houming, et al.
Veröffentlicht: (2025)
von: Qiu, Houming, et al.
Veröffentlicht: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
von: Woo, Hyein, et al.
Veröffentlicht: (2025)
von: Woo, Hyein, et al.
Veröffentlicht: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
von: Li, Youjie, et al.
Veröffentlicht: (2025) -
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
von: Zhang, Ruisi, et al.
Veröffentlicht: (2024) -
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025) -
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
von: Ma, Qianli, et al.
Veröffentlicht: (2025) -
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)