Understanding and Improving Communication Performance in Multi-node LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Singhania, Prajwal, Singh, Siddharth, Hough, Lannie Dalton, Srivastava, Akarsh, Menon, Harshitha, Jekel, Charles Fredrick, Bhatele, Abhinav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimizing Agentic Language Model Inference via Speculative Tool Calls
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
di: Singh, Siddharth, et al.
Pubblicazione: (2023)
di: Singh, Siddharth, et al.
Pubblicazione: (2023)
HPC-Coder: Modeling Parallel Programs using Large Language Models
di: Nichols, Daniel, et al.
Pubblicazione: (2023)
di: Nichols, Daniel, et al.
Pubblicazione: (2023)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
di: Singh, Siddharth, et al.
Pubblicazione: (2025)
di: Singh, Siddharth, et al.
Pubblicazione: (2025)
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
Performance-Aligned LLMs for Generating Fast Code
di: Nichols, Daniel, et al.
Pubblicazione: (2024)
di: Nichols, Daniel, et al.
Pubblicazione: (2024)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
di: Chaturvedi, Aman, et al.
Pubblicazione: (2024)
di: Chaturvedi, Aman, et al.
Pubblicazione: (2024)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
di: Ranjan, Aditya K., et al.
Pubblicazione: (2025)
di: Ranjan, Aditya K., et al.
Pubblicazione: (2025)
The Big Send-off: Scalable and Performant Collectives for Deep Learning
di: Singh, Siddharth, et al.
Pubblicazione: (2025)
di: Singh, Siddharth, et al.
Pubblicazione: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
di: Cankur, Onur, et al.
Pubblicazione: (2025)
di: Cankur, Onur, et al.
Pubblicazione: (2025)
ParEval-Repo: A Benchmark Suite for Evaluating LLMs with Repository-level HPC Translation Tasks
di: Davis, Joshua H., et al.
Pubblicazione: (2025)
di: Davis, Joshua H., et al.
Pubblicazione: (2025)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
di: Xu, Yiheng, et al.
Pubblicazione: (2023)
di: Xu, Yiheng, et al.
Pubblicazione: (2023)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
di: Costello, Ian J., et al.
Pubblicazione: (2020)
di: Costello, Ian J., et al.
Pubblicazione: (2020)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
Pipit: Scripting the analysis of parallel execution traces
di: Bhatele, Abhinav, et al.
Pubblicazione: (2023)
di: Bhatele, Abhinav, et al.
Pubblicazione: (2023)
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
di: Fink, Zane, et al.
Pubblicazione: (2024)
di: Fink, Zane, et al.
Pubblicazione: (2024)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
di: Arya, Mayank, et al.
Pubblicazione: (2025)
di: Arya, Mayank, et al.
Pubblicazione: (2025)
Federated Inference for Heterogeneous LLM Communication and Collaboration
di: Chen, Zihan, et al.
Pubblicazione: (2026)
di: Chen, Zihan, et al.
Pubblicazione: (2026)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
Automated Programmatic Performance Analysis of Parallel Programs
di: Cankur, Onur, et al.
Pubblicazione: (2024)
di: Cankur, Onur, et al.
Pubblicazione: (2024)
Can Large Language Models Write Parallel Code?
di: Nichols, Daniel, et al.
Pubblicazione: (2024)
di: Nichols, Daniel, et al.
Pubblicazione: (2024)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
di: Wei, Cunyang, et al.
Pubblicazione: (2026)
di: Wei, Cunyang, et al.
Pubblicazione: (2026)
LLM-assisted Agentic Edge Intelligence Framework
di: Dehury, Chinmaya Kumar, et al.
Pubblicazione: (2026)
di: Dehury, Chinmaya Kumar, et al.
Pubblicazione: (2026)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
di: Wang, Weiye, et al.
Pubblicazione: (2026)
di: Wang, Weiye, et al.
Pubblicazione: (2026)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
di: Zhang, Songge, et al.
Pubblicazione: (2026)
di: Zhang, Songge, et al.
Pubblicazione: (2026)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
di: Chen, Jiu, et al.
Pubblicazione: (2026)
di: Chen, Jiu, et al.
Pubblicazione: (2026)
Pandemics In Silico: Scaling an Agent-Based Simulation on Realistic Social Contact Networks
di: Kitson, Joy, et al.
Pubblicazione: (2024)
di: Kitson, Joy, et al.
Pubblicazione: (2024)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
di: Wang, Rongxiang, et al.
Pubblicazione: (2025)
di: Wang, Rongxiang, et al.
Pubblicazione: (2025)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025)
di: Wu, Panlong, et al.
Pubblicazione: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
di: Zhang, Kai, et al.
Pubblicazione: (2025)
di: Zhang, Kai, et al.
Pubblicazione: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
WANSpec: Leveraging Global Compute Capacity for LLM Inference
di: Martin, Noah, et al.
Pubblicazione: (2026)
di: Martin, Noah, et al.
Pubblicazione: (2026)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Optimizing Agentic Language Model Inference via Speculative Tool Calls
di: Nichols, Daniel, et al.
Pubblicazione: (2025) -
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
di: Nichols, Daniel, et al.
Pubblicazione: (2025) -
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
di: Singh, Siddharth, et al.
Pubblicazione: (2023) -
HPC-Coder: Modeling Parallel Programs using Large Language Models
di: Nichols, Daniel, et al.
Pubblicazione: (2023) -
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
di: Singh, Siddharth, et al.
Pubblicazione: (2025)