Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
Fuente:
arXiv
Guardado en:
| Autores principales: | Prieto, Pablo, Abad, Pablo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
por: Yin, Wangsong, et al.
Publicado: (2025)
por: Yin, Wangsong, et al.
Publicado: (2025)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
por: Yao, Feiyu, et al.
Publicado: (2026)
por: Yao, Feiyu, et al.
Publicado: (2026)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
Deploying Open-Source Large Language Models: A performance Analysis
por: Bendi-Ouis, Yannis, et al.
Publicado: (2024)
por: Bendi-Ouis, Yannis, et al.
Publicado: (2024)
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
por: Paramanayakam, Varatheepan, et al.
Publicado: (2025)
por: Paramanayakam, Varatheepan, et al.
Publicado: (2025)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
por: Almurshed, Osama, et al.
Publicado: (2025)
por: Almurshed, Osama, et al.
Publicado: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
por: Chai, Yuji, et al.
Publicado: (2025)
por: Chai, Yuji, et al.
Publicado: (2025)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
por: Zhao, Youpeng, et al.
Publicado: (2024)
por: Zhao, Youpeng, et al.
Publicado: (2024)
On the Sustainability of AI Inferences in the Edge
por: Sobhani, Ghazal, et al.
Publicado: (2025)
por: Sobhani, Ghazal, et al.
Publicado: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
por: Knoop, Jonathan, et al.
Publicado: (2026)
por: Knoop, Jonathan, et al.
Publicado: (2026)
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
por: Pinnock, Alyssa, et al.
Publicado: (2025)
por: Pinnock, Alyssa, et al.
Publicado: (2025)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
por: Zhang, Tuo, et al.
Publicado: (2025)
por: Zhang, Tuo, et al.
Publicado: (2025)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
por: Vishwanathan, Manoj, et al.
Publicado: (2026)
por: Vishwanathan, Manoj, et al.
Publicado: (2026)
Fairness in Serving Large Language Models
por: Sheng, Ying, et al.
Publicado: (2023)
por: Sheng, Ying, et al.
Publicado: (2023)
KernelBench: Can LLMs Write Efficient GPU Kernels?
por: Ouyang, Anne, et al.
Publicado: (2025)
por: Ouyang, Anne, et al.
Publicado: (2025)
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
por: Jang, SiYoung, et al.
Publicado: (2025)
por: Jang, SiYoung, et al.
Publicado: (2025)
SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2026)
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2026)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
por: Andrews, Martin, et al.
Publicado: (2025)
por: Andrews, Martin, et al.
Publicado: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
por: Hossain, Md Arafat, et al.
Publicado: (2025)
por: Hossain, Md Arafat, et al.
Publicado: (2025)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
por: d'Aloisio, Giordano, et al.
Publicado: (2024)
por: d'Aloisio, Giordano, et al.
Publicado: (2024)
Personalized Model-Based Design of Human Centric AI enabled CPS for Long term usage
por: Ngabonziza, Bernard, et al.
Publicado: (2026)
por: Ngabonziza, Bernard, et al.
Publicado: (2026)
Breaking the Loop: Detecting and Mitigating Denial-of-Service Vulnerabilities in Large Language Models
por: Yu, Junzhe, et al.
Publicado: (2025)
por: Yu, Junzhe, et al.
Publicado: (2025)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
por: Liargkovas, Georgios, et al.
Publicado: (2026)
por: Liargkovas, Georgios, et al.
Publicado: (2026)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
por: Abstreiter, Maximilian, et al.
Publicado: (2025)
por: Abstreiter, Maximilian, et al.
Publicado: (2025)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
por: Ma, Jeffrey Jian, et al.
Publicado: (2025)
por: Ma, Jeffrey Jian, et al.
Publicado: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
por: Zhao, Youpeng, et al.
Publicado: (2024)
por: Zhao, Youpeng, et al.
Publicado: (2024)
Benchmark-based Study of CPU/GPU Power-Related Features through JAX and TensorFlow
por: Tchakoute, Roblex Nana, et al.
Publicado: (2025)
por: Tchakoute, Roblex Nana, et al.
Publicado: (2025)
Revealing NVIDIA Closed-Source Driver Command Streams for CPU-GPU Runtime Behavior Insight
por: Yan, Yuang, et al.
Publicado: (2026)
por: Yan, Yuang, et al.
Publicado: (2026)
Reliability by design: quantifying and eliminating fabrication risk in LLMs. From generative to consultative AI: a comparative analysis in the legal domain and lessons for high-stakes knowledge bases
por: Dantart, Alex
Publicado: (2026)
por: Dantart, Alex
Publicado: (2026)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
por: Zhao, Yushang, et al.
Publicado: (2025)
por: Zhao, Yushang, et al.
Publicado: (2025)
Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models
por: Rosas, Miguel Romero, et al.
Publicado: (2024)
por: Rosas, Miguel Romero, et al.
Publicado: (2024)
Offloading and Quality Control for AI Generated Content Services in 6G Mobile Edge Computing Networks
por: Wang, Yitong, et al.
Publicado: (2023)
por: Wang, Yitong, et al.
Publicado: (2023)
MambaCPU: Enhanced Correlation Mining with State Space Models for CPU Performance Prediction
por: Liu, Xiaoman
Publicado: (2024)
por: Liu, Xiaoman
Publicado: (2024)
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
por: Arantes, Gabriel M., et al.
Publicado: (2025)
por: Arantes, Gabriel M., et al.
Publicado: (2025)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
por: Ivanov, Andrei, et al.
Publicado: (2025)
por: Ivanov, Andrei, et al.
Publicado: (2025)
PixLift: Accelerating Web Browsing via AI Upscaling
por: Atinafu, Yonas, et al.
Publicado: (2025)
por: Atinafu, Yonas, et al.
Publicado: (2025)
Tiny-QMoE
por: Cashman, Jack, et al.
Publicado: (2025)
por: Cashman, Jack, et al.
Publicado: (2025)
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
por: Pfister, Rolf, et al.
Publicado: (2025)
por: Pfister, Rolf, et al.
Publicado: (2025)
XTC, A Research Platform for Optimizing AI Workload Operators
por: Hugo, Pompougnac, et al.
Publicado: (2025)
por: Hugo, Pompougnac, et al.
Publicado: (2025)
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
por: Cai, Yanan, et al.
Publicado: (2025)
por: Cai, Yanan, et al.
Publicado: (2025)
Ejemplares similares
-
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
por: Yin, Wangsong, et al.
Publicado: (2025) -
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
por: Yao, Feiyu, et al.
Publicado: (2026) -
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
por: Vellaisamy, Prabhu, et al.
Publicado: (2025) -
Deploying Open-Source Large Language Models: A performance Analysis
por: Bendi-Ouis, Yannis, et al.
Publicado: (2024) -
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
por: Paramanayakam, Varatheepan, et al.
Publicado: (2025)