LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
Fuente:
arXiv
Guardado en:
| Autores principales: | Tummalapalli, Pranay, Arayakandy, Sahil, Pal, Ritam, Kundan, Kautuk |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
por: Hao, Zixu, et al.
Publicado: (2025)
por: Hao, Zixu, et al.
Publicado: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
por: Rajashekar, Kolichala, et al.
Publicado: (2025)
por: Rajashekar, Kolichala, et al.
Publicado: (2025)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
por: Dai, Yuntao, et al.
Publicado: (2026)
por: Dai, Yuntao, et al.
Publicado: (2026)
AeroGen: Agentic Drone Autonomy through Single-Shot Structured Prompting & Drone SDK
por: Astu, Kautuk, et al.
Publicado: (2026)
por: Astu, Kautuk, et al.
Publicado: (2026)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
por: Lin, Shouxu, et al.
Publicado: (2026)
por: Lin, Shouxu, et al.
Publicado: (2026)
Edge AI in Highly Volatile Environments: Is Fairness Worth the Accuracy Trade-off?
por: Zaland, Obaidullah, et al.
Publicado: (2025)
por: Zaland, Obaidullah, et al.
Publicado: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
por: Arya, Mayank, et al.
Publicado: (2025)
por: Arya, Mayank, et al.
Publicado: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
por: Zhang, Haolin, et al.
Publicado: (2025)
por: Zhang, Haolin, et al.
Publicado: (2025)
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
por: Malekzadeh, Mohammad, et al.
Publicado: (2023)
por: Malekzadeh, Mohammad, et al.
Publicado: (2023)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
por: Recasens, Pol G., et al.
Publicado: (2025)
por: Recasens, Pol G., et al.
Publicado: (2025)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
por: Nicolae, Radu, et al.
Publicado: (2026)
por: Nicolae, Radu, et al.
Publicado: (2026)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
por: Li, Zhuojin, et al.
Publicado: (2025)
por: Li, Zhuojin, et al.
Publicado: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
por: Arif, Moiz, et al.
Publicado: (2026)
por: Arif, Moiz, et al.
Publicado: (2026)
AeroDaaS: A Programmable Drones-as-a-Service Platform for Intelligent Aerial Systems
por: Astu, Kautuk, et al.
Publicado: (2026)
por: Astu, Kautuk, et al.
Publicado: (2026)
AeroDaaS: Towards an Application Programming Framework for Drones-as-a-Service
por: Raj, Suman, et al.
Publicado: (2025)
por: Raj, Suman, et al.
Publicado: (2025)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
por: Guo, Runsheng Benson, et al.
Publicado: (2025)
por: Guo, Runsheng Benson, et al.
Publicado: (2025)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
por: Abstreiter, Maximilian, et al.
Publicado: (2025)
por: Abstreiter, Maximilian, et al.
Publicado: (2025)
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
por: Chen, Huamin, et al.
Publicado: (2026)
por: Chen, Huamin, et al.
Publicado: (2026)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
por: Maczan, Jędrzej
Publicado: (2026)
por: Maczan, Jędrzej
Publicado: (2026)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
por: Gao, Lei, et al.
Publicado: (2024)
por: Gao, Lei, et al.
Publicado: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
por: Mo, Zizhao, et al.
Publicado: (2026)
por: Mo, Zizhao, et al.
Publicado: (2026)
WindVE: Collaborative CPU-NPU Vector Embedding
por: Huang, Jinqi, et al.
Publicado: (2025)
por: Huang, Jinqi, et al.
Publicado: (2025)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
por: Levine, Reese, et al.
Publicado: (2026)
por: Levine, Reese, et al.
Publicado: (2026)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
por: Zhang, Hongbin, et al.
Publicado: (2026)
por: Zhang, Hongbin, et al.
Publicado: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
por: Zhang, Mingjin, et al.
Publicado: (2024)
por: Zhang, Mingjin, et al.
Publicado: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
por: Singhania, Prajwal, et al.
Publicado: (2025)
por: Singhania, Prajwal, et al.
Publicado: (2025)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
por: Zhang, Tianyi, et al.
Publicado: (2025)
por: Zhang, Tianyi, et al.
Publicado: (2025)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
por: Li, Jinhao, et al.
Publicado: (2023)
por: Li, Jinhao, et al.
Publicado: (2023)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
por: Fang, Jingzhi, et al.
Publicado: (2025)
por: Fang, Jingzhi, et al.
Publicado: (2025)
GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters
por: Esposito, Giuseppe, et al.
Publicado: (2025)
por: Esposito, Giuseppe, et al.
Publicado: (2025)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
por: Wang, Chong, et al.
Publicado: (2026)
por: Wang, Chong, et al.
Publicado: (2026)
Decentralized Task Offloading and Load-Balancing for Mobile Edge Computing in Dense Networks
por: Yahya, Mariam, et al.
Publicado: (2024)
por: Yahya, Mariam, et al.
Publicado: (2024)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
por: Bournias, Ilias, et al.
Publicado: (2024)
por: Bournias, Ilias, et al.
Publicado: (2024)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
por: Tian, Jian, et al.
Publicado: (2025)
por: Tian, Jian, et al.
Publicado: (2025)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
por: Shu, Zhihao, et al.
Publicado: (2026)
por: Shu, Zhihao, et al.
Publicado: (2026)
MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
por: Moore, H., et al.
Publicado: (2026)
por: Moore, H., et al.
Publicado: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
por: Fan, Jiakun, et al.
Publicado: (2025)
por: Fan, Jiakun, et al.
Publicado: (2025)
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
por: Mukherjee, Soutrik, et al.
Publicado: (2026)
por: Mukherjee, Soutrik, et al.
Publicado: (2026)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
por: Wang, Liujianfu, et al.
Publicado: (2025)
por: Wang, Liujianfu, et al.
Publicado: (2025)
Peformance Isolation for Inference Processes in Edge GPU Systems
por: Martín, Juan José, et al.
Publicado: (2026)
por: Martín, Juan José, et al.
Publicado: (2026)
Ejemplares similares
-
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
por: Hao, Zixu, et al.
Publicado: (2025) -
Toward Sustainability-Aware LLM Inference on Edge Clusters
por: Rajashekar, Kolichala, et al.
Publicado: (2025) -
Accelerating OpenPangu Inference on NPU via Speculative Decoding
por: Dai, Yuntao, et al.
Publicado: (2026) -
AeroGen: Agentic Drone Autonomy through Single-Shot Structured Prompting & Drone SDK
por: Astu, Kautuk, et al.
Publicado: (2026) -
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
por: Lin, Shouxu, et al.
Publicado: (2026)