Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Trifan, Octavian Alexandru, Sangaiah, Karthik, Awad, Muhammad, Osama, Muhammad, Gudaparthi, Sumanth, Nicolau, Alexandru, Veidenbaum, Alexander, Dasika, Ganesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
Seer: Predictive Runtime Kernel Selection for Irregular Problems
von: Swann, Ryan, et al.
Veröffentlicht: (2024)
von: Swann, Ryan, et al.
Veröffentlicht: (2024)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
Iris: First-Class Multi-GPU Programming Experience in Triton
von: Awad, Muhammad, et al.
Veröffentlicht: (2025)
von: Awad, Muhammad, et al.
Veröffentlicht: (2025)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
von: Swann, Ryan, et al.
Veröffentlicht: (2025)
von: Swann, Ryan, et al.
Veröffentlicht: (2025)
Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
von: Bouvier, Thomas, et al.
Veröffentlicht: (2024)
von: Bouvier, Thomas, et al.
Veröffentlicht: (2024)
Cppless: Single-Source and High-Performance Serverless Programming in C++
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
In Serverless, OS Scheduler Choice Costs Money: A Hybrid Scheduling Approach for Cheaper FaaS
von: Zhao, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhao, Yuxuan, et al.
Veröffentlicht: (2024)
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
Leveraging Mathematical Reasoning of LLMs for Efficient GPU Thread Mapping
von: Maureira, Jose, et al.
Veröffentlicht: (2026)
von: Maureira, Jose, et al.
Veröffentlicht: (2026)
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
SIMT/GPU Data Race Verification using ISCC and Intermediary Code Representations: A Case Study
von: Osterhout, Andrew, et al.
Veröffentlicht: (2025)
von: Osterhout, Andrew, et al.
Veröffentlicht: (2025)
HiRace: Accurate and Fast Source-Level Race Checking of GPU Programs
von: Jacobson, John, et al.
Veröffentlicht: (2024)
von: Jacobson, John, et al.
Veröffentlicht: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
Channel Prediction under Network Distribution Shift Using Continual Learning-based Loss Regularization
von: Mohsin, Muhammad Ahmed, et al.
Veröffentlicht: (2025)
von: Mohsin, Muhammad Ahmed, et al.
Veröffentlicht: (2025)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026)
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
Distributed Inference Performance Optimization for LLMs on CPUs
von: He, Pujiang, et al.
Veröffentlicht: (2024)
von: He, Pujiang, et al.
Veröffentlicht: (2024)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
von: Copik, Marcin, et al.
Veröffentlicht: (2022)
von: Copik, Marcin, et al.
Veröffentlicht: (2022)
Software Resource Disaggregation for HPC with Serverless Computing
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
Communication Efficient Byzantine Agreement with Predictions
von: Dzulfikar, Muhammad Ayaz, et al.
Veröffentlicht: (2026)
von: Dzulfikar, Muhammad Ayaz, et al.
Veröffentlicht: (2026)
Exploring Novel Data Storage Approaches for Large-Scale Numerical Weather Prediction
von: Gil, Nicolau Manubens
Veröffentlicht: (2026)
von: Gil, Nicolau Manubens
Veröffentlicht: (2026)
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
von: V., Karthik, et al.
Veröffentlicht: (2024)
von: V., Karthik, et al.
Veröffentlicht: (2024)
A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates
von: Mao, Zirui, et al.
Veröffentlicht: (2023)
von: Mao, Zirui, et al.
Veröffentlicht: (2023)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
von: Neff, Reece, et al.
Veröffentlicht: (2023)
von: Neff, Reece, et al.
Veröffentlicht: (2023)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
von: Xu, Zhihao, et al.
Veröffentlicht: (2025)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
von: Iwabuchi, Keita, et al.
Veröffentlicht: (2026)
von: Iwabuchi, Keita, et al.
Veröffentlicht: (2026)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
von: Yang, Zhuoping, et al.
Veröffentlicht: (2025)
von: Yang, Zhuoping, et al.
Veröffentlicht: (2025)
Efficient Accelerated Graph Edit Distance Computation on GPU
von: Dabah, Adel, et al.
Veröffentlicht: (2026)
von: Dabah, Adel, et al.
Veröffentlicht: (2026)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025)
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
FastTrack: GPU-Accelerated Tracking for Visual SLAM
von: Khabiri, Kimia, et al.
Veröffentlicht: (2025)
von: Khabiri, Kimia, et al.
Veröffentlicht: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025) -
Seer: Predictive Runtime Kernel Selection for Irregular Problems
von: Swann, Ryan, et al.
Veröffentlicht: (2024) -
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
von: Tschand, Arya, et al.
Veröffentlicht: (2025) -
Iris: First-Class Multi-GPU Programming Experience in Triton
von: Awad, Muhammad, et al.
Veröffentlicht: (2025) -
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
von: Swann, Ryan, et al.
Veröffentlicht: (2025)