Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Arif, Moiz, Maurya, Avinash, Vazhkudai, Sudharshan, Nicolae, Bogdan |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
par: Maurya, Avinash, et autres
Publié: (2026)
par: Maurya, Avinash, et autres
Publié: (2026)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
par: Gossman, Mikaila J., et autres
Publié: (2025)
par: Gossman, Mikaila J., et autres
Publié: (2025)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
par: Maurya, Avinash, et autres
Publié: (2024)
par: Maurya, Avinash, et autres
Publié: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
par: Karfakis, George, et autres
Publié: (2025)
par: Karfakis, George, et autres
Publié: (2025)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
par: Grbic, Dragana
Publié: (2026)
par: Grbic, Dragana
Publié: (2026)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
par: Maurya, Avinash, et autres
Publié: (2024)
par: Maurya, Avinash, et autres
Publié: (2024)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
par: Villalobos, Johansell, et autres
Publié: (2025)
par: Villalobos, Johansell, et autres
Publié: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
par: Boudaoud, Afif, et autres
Publié: (2026)
par: Boudaoud, Afif, et autres
Publié: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
par: Zhang, Li, et autres
Publié: (2025)
par: Zhang, Li, et autres
Publié: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
par: Proaño, Andrès Rubio, et autres
Publié: (2024)
par: Proaño, Andrès Rubio, et autres
Publié: (2024)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
par: Ng, Nathan, et autres
Publié: (2026)
par: Ng, Nathan, et autres
Publié: (2026)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
par: Dutt, Anurag, et autres
Publié: (2025)
par: Dutt, Anurag, et autres
Publié: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
par: Zhao, Yanbo, et autres
Publié: (2025)
par: Zhao, Yanbo, et autres
Publié: (2025)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
par: Debnath, Shimul, et autres
Publié: (2026)
par: Debnath, Shimul, et autres
Publié: (2026)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
par: Lin, Mao, et autres
Publié: (2026)
par: Lin, Mao, et autres
Publié: (2026)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
par: Zhao, Xuanlei, et autres
Publié: (2024)
par: Zhao, Xuanlei, et autres
Publié: (2024)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
par: Islam, Tanzima Z., et autres
Publié: (2024)
par: Islam, Tanzima Z., et autres
Publié: (2024)
PICO: Performance Insights for Collective Operations
par: Pasqualoni, Saverio, et autres
Publié: (2025)
par: Pasqualoni, Saverio, et autres
Publié: (2025)
A Performance Analysis of BFT Consensus for Blockchains
par: Chan, J. D., et autres
Publié: (2024)
par: Chan, J. D., et autres
Publié: (2024)
Scalable GPU Performance Variability Analysis framework
par: Lahiry, Ankur, et autres
Publié: (2025)
par: Lahiry, Ankur, et autres
Publié: (2025)
Automated Programmatic Performance Analysis of Parallel Programs
par: Cankur, Onur, et autres
Publié: (2024)
par: Cankur, Onur, et autres
Publié: (2024)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
par: Zhang, Yaozheng, et autres
Publié: (2025)
par: Zhang, Yaozheng, et autres
Publié: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
par: Vatsavai, Sairam Sri, et autres
Publié: (2025)
par: Vatsavai, Sairam Sri, et autres
Publié: (2025)
Performance Debugging through Microarchitectural Sensitivity and Causality Analysis
par: Dutilleul, Alban, et autres
Publié: (2024)
par: Dutilleul, Alban, et autres
Publié: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
par: de Morais, Gustavo, et autres
Publié: (2025)
par: de Morais, Gustavo, et autres
Publié: (2025)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
par: Aqasizade, Hossein, et autres
Publié: (2024)
par: Aqasizade, Hossein, et autres
Publié: (2024)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
par: Banegas-Luna, Antonio Jesús, et autres
Publié: (2025)
par: Banegas-Luna, Antonio Jesús, et autres
Publié: (2025)
Taking GPU Programming Models to Task for Performance Portability
par: Davis, Joshua H., et autres
Publié: (2024)
par: Davis, Joshua H., et autres
Publié: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
par: Zhuang, Chen, et autres
Publié: (2024)
par: Zhuang, Chen, et autres
Publié: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
par: Wang, Yuxin, et autres
Publié: (2023)
par: Wang, Yuxin, et autres
Publié: (2023)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
par: Xu, Jingwei, et autres
Publié: (2025)
par: Xu, Jingwei, et autres
Publié: (2025)
Mitigating GIL Bottlenecks in Edge AI Systems
par: Mandal, Mridankan, et autres
Publié: (2026)
par: Mandal, Mridankan, et autres
Publié: (2026)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
par: Davis, Joshua H., et autres
Publié: (2026)
par: Davis, Joshua H., et autres
Publié: (2026)
eBPF-Based Instrumentation for Generalisable Diagnosis of Performance Degradation
par: Landau, Diogo, et autres
Publié: (2025)
par: Landau, Diogo, et autres
Publié: (2025)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
par: Abstreiter, Maximilian, et autres
Publié: (2025)
par: Abstreiter, Maximilian, et autres
Publié: (2025)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
par: Suffa, Philipp, et autres
Publié: (2024)
par: Suffa, Philipp, et autres
Publié: (2024)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
par: Pilliat, Emmanuel
Publié: (2026)
par: Pilliat, Emmanuel
Publié: (2026)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
par: Kulkarni, Apurv Deepak, et autres
Publié: (2025)
par: Kulkarni, Apurv Deepak, et autres
Publié: (2025)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
par: Afzal, Ayesha, et autres
Publié: (2025)
par: Afzal, Ayesha, et autres
Publié: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
par: Ramesh, Risshab Srinivas
Publié: (2024)
par: Ramesh, Risshab Srinivas
Publié: (2024)
Documents similaires
-
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
par: Maurya, Avinash, et autres
Publié: (2026) -
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
par: Gossman, Mikaila J., et autres
Publié: (2025) -
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
par: Maurya, Avinash, et autres
Publié: (2024) -
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
par: Karfakis, George, et autres
Publié: (2025) -
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
par: Grbic, Dragana
Publié: (2026)