Understanding and Improving Communication Performance in Multi-node LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singhania, Prajwal, Singh, Siddharth, Hough, Lannie Dalton, Srivastava, Akarsh, Menon, Harshitha, Jekel, Charles Fredrick, Bhatele, Abhinav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing Agentic Language Model Inference via Speculative Tool Calls
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
HPC-Coder: Modeling Parallel Programs using Large Language Models
von: Nichols, Daniel, et al.
Veröffentlicht: (2023)
von: Nichols, Daniel, et al.
Veröffentlicht: (2023)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
Taking GPU Programming Models to Task for Performance Portability
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
Performance-Aligned LLMs for Generating Fast Code
von: Nichols, Daniel, et al.
Veröffentlicht: (2024)
von: Nichols, Daniel, et al.
Veröffentlicht: (2024)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
von: Chaturvedi, Aman, et al.
Veröffentlicht: (2024)
von: Chaturvedi, Aman, et al.
Veröffentlicht: (2024)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
von: Ranjan, Aditya K., et al.
Veröffentlicht: (2025)
von: Ranjan, Aditya K., et al.
Veröffentlicht: (2025)
The Big Send-off: Scalable and Performant Collectives for Deep Learning
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
von: Cankur, Onur, et al.
Veröffentlicht: (2025)
ParEval-Repo: A Benchmark Suite for Evaluating LLMs with Repository-level HPC Translation Tasks
von: Davis, Joshua H., et al.
Veröffentlicht: (2025)
von: Davis, Joshua H., et al.
Veröffentlicht: (2025)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
von: Xu, Yiheng, et al.
Veröffentlicht: (2023)
von: Xu, Yiheng, et al.
Veröffentlicht: (2023)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
von: Costello, Ian J., et al.
Veröffentlicht: (2020)
von: Costello, Ian J., et al.
Veröffentlicht: (2020)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
Pipit: Scripting the analysis of parallel execution traces
von: Bhatele, Abhinav, et al.
Veröffentlicht: (2023)
von: Bhatele, Abhinav, et al.
Veröffentlicht: (2023)
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
von: Fink, Zane, et al.
Veröffentlicht: (2024)
von: Fink, Zane, et al.
Veröffentlicht: (2024)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
Federated Inference for Heterogeneous LLM Communication and Collaboration
von: Chen, Zihan, et al.
Veröffentlicht: (2026)
von: Chen, Zihan, et al.
Veröffentlicht: (2026)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
Automated Programmatic Performance Analysis of Parallel Programs
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
Can Large Language Models Write Parallel Code?
von: Nichols, Daniel, et al.
Veröffentlicht: (2024)
von: Nichols, Daniel, et al.
Veröffentlicht: (2024)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
von: Wei, Cunyang, et al.
Veröffentlicht: (2026)
von: Wei, Cunyang, et al.
Veröffentlicht: (2026)
LLM-assisted Agentic Edge Intelligence Framework
von: Dehury, Chinmaya Kumar, et al.
Veröffentlicht: (2026)
von: Dehury, Chinmaya Kumar, et al.
Veröffentlicht: (2026)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
von: Zhang, Songge, et al.
Veröffentlicht: (2026)
von: Zhang, Songge, et al.
Veröffentlicht: (2026)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
von: Chen, Jiu, et al.
Veröffentlicht: (2026)
von: Chen, Jiu, et al.
Veröffentlicht: (2026)
Pandemics In Silico: Scaling an Agent-Based Simulation on Realistic Social Contact Networks
von: Kitson, Joy, et al.
Veröffentlicht: (2024)
von: Kitson, Joy, et al.
Veröffentlicht: (2024)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2025)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
WANSpec: Leveraging Global Compute Capacity for LLM Inference
von: Martin, Noah, et al.
Veröffentlicht: (2026)
von: Martin, Noah, et al.
Veröffentlicht: (2026)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Optimizing Agentic Language Model Inference via Speculative Tool Calls
von: Nichols, Daniel, et al.
Veröffentlicht: (2025) -
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
von: Nichols, Daniel, et al.
Veröffentlicht: (2025) -
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
von: Singh, Siddharth, et al.
Veröffentlicht: (2023) -
HPC-Coder: Modeling Parallel Programs using Large Language Models
von: Nichols, Daniel, et al.
Veröffentlicht: (2023) -
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)