LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tummalapalli, Pranay, Arayakandy, Sahil, Pal, Ritam, Kundan, Kautuk |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
par: Hao, Zixu, et autres
Publié: (2025)
par: Hao, Zixu, et autres
Publié: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
par: Rajashekar, Kolichala, et autres
Publié: (2025)
par: Rajashekar, Kolichala, et autres
Publié: (2025)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
par: Dai, Yuntao, et autres
Publié: (2026)
par: Dai, Yuntao, et autres
Publié: (2026)
AeroGen: Agentic Drone Autonomy through Single-Shot Structured Prompting & Drone SDK
par: Astu, Kautuk, et autres
Publié: (2026)
par: Astu, Kautuk, et autres
Publié: (2026)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
par: Lin, Shouxu, et autres
Publié: (2026)
par: Lin, Shouxu, et autres
Publié: (2026)
Edge AI in Highly Volatile Environments: Is Fairness Worth the Accuracy Trade-off?
par: Zaland, Obaidullah, et autres
Publié: (2025)
par: Zaland, Obaidullah, et autres
Publié: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
par: Arya, Mayank, et autres
Publié: (2025)
par: Arya, Mayank, et autres
Publié: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
par: Zhang, Haolin, et autres
Publié: (2025)
par: Zhang, Haolin, et autres
Publié: (2025)
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
par: Malekzadeh, Mohammad, et autres
Publié: (2023)
par: Malekzadeh, Mohammad, et autres
Publié: (2023)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
par: Recasens, Pol G., et autres
Publié: (2025)
par: Recasens, Pol G., et autres
Publié: (2025)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
par: Nicolae, Radu, et autres
Publié: (2026)
par: Nicolae, Radu, et autres
Publié: (2026)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
par: Li, Zhuojin, et autres
Publié: (2025)
par: Li, Zhuojin, et autres
Publié: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
par: Arif, Moiz, et autres
Publié: (2026)
par: Arif, Moiz, et autres
Publié: (2026)
AeroDaaS: A Programmable Drones-as-a-Service Platform for Intelligent Aerial Systems
par: Astu, Kautuk, et autres
Publié: (2026)
par: Astu, Kautuk, et autres
Publié: (2026)
AeroDaaS: Towards an Application Programming Framework for Drones-as-a-Service
par: Raj, Suman, et autres
Publié: (2025)
par: Raj, Suman, et autres
Publié: (2025)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
par: Guo, Runsheng Benson, et autres
Publié: (2025)
par: Guo, Runsheng Benson, et autres
Publié: (2025)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
par: Abstreiter, Maximilian, et autres
Publié: (2025)
par: Abstreiter, Maximilian, et autres
Publié: (2025)
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
par: Maczan, Jędrzej
Publié: (2026)
par: Maczan, Jędrzej
Publié: (2026)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
par: Gao, Lei, et autres
Publié: (2024)
par: Gao, Lei, et autres
Publié: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026)
par: Mo, Zizhao, et autres
Publié: (2026)
WindVE: Collaborative CPU-NPU Vector Embedding
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
par: Levine, Reese, et autres
Publié: (2026)
par: Levine, Reese, et autres
Publié: (2026)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
par: Zhang, Hongbin, et autres
Publié: (2026)
par: Zhang, Hongbin, et autres
Publié: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
par: Zhang, Mingjin, et autres
Publié: (2024)
par: Zhang, Mingjin, et autres
Publié: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
par: Singhania, Prajwal, et autres
Publié: (2025)
par: Singhania, Prajwal, et autres
Publié: (2025)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
par: Zhang, Tianyi, et autres
Publié: (2025)
par: Zhang, Tianyi, et autres
Publié: (2025)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
par: Li, Jinhao, et autres
Publié: (2023)
par: Li, Jinhao, et autres
Publié: (2023)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
par: Fang, Jingzhi, et autres
Publié: (2025)
par: Fang, Jingzhi, et autres
Publié: (2025)
GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters
par: Esposito, Giuseppe, et autres
Publié: (2025)
par: Esposito, Giuseppe, et autres
Publié: (2025)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
par: Wang, Chong, et autres
Publié: (2026)
par: Wang, Chong, et autres
Publié: (2026)
Decentralized Task Offloading and Load-Balancing for Mobile Edge Computing in Dense Networks
par: Yahya, Mariam, et autres
Publié: (2024)
par: Yahya, Mariam, et autres
Publié: (2024)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
par: Bournias, Ilias, et autres
Publié: (2024)
par: Bournias, Ilias, et autres
Publié: (2024)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
par: Tian, Jian, et autres
Publié: (2025)
par: Tian, Jian, et autres
Publié: (2025)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
par: Shu, Zhihao, et autres
Publié: (2026)
par: Shu, Zhihao, et autres
Publié: (2026)
MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
par: Moore, H., et autres
Publié: (2026)
par: Moore, H., et autres
Publié: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
par: Fan, Jiakun, et autres
Publié: (2025)
par: Fan, Jiakun, et autres
Publié: (2025)
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
par: Mukherjee, Soutrik, et autres
Publié: (2026)
par: Mukherjee, Soutrik, et autres
Publié: (2026)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
par: Wang, Liujianfu, et autres
Publié: (2025)
par: Wang, Liujianfu, et autres
Publié: (2025)
Peformance Isolation for Inference Processes in Edge GPU Systems
par: Martín, Juan José, et autres
Publié: (2026)
par: Martín, Juan José, et autres
Publié: (2026)
Documents similaires
-
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
par: Hao, Zixu, et autres
Publié: (2025) -
Toward Sustainability-Aware LLM Inference on Edge Clusters
par: Rajashekar, Kolichala, et autres
Publié: (2025) -
Accelerating OpenPangu Inference on NPU via Speculative Decoding
par: Dai, Yuntao, et autres
Publié: (2026) -
AeroGen: Agentic Drone Autonomy through Single-Shot Structured Prompting & Drone SDK
par: Astu, Kautuk, et autres
Publié: (2026) -
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
par: Lin, Shouxu, et autres
Publié: (2026)