Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Renney, Harri, Trad, Fouad, Mattarock, Michael, Wood, Zena
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915961903251456
author Renney, Harri
Trad, Fouad
Mattarock, Michael
Wood, Zena
author_facet Renney, Harri
Trad, Fouad
Mattarock, Michael
Wood, Zena
contents Large language models (LLMs) are becoming increasingly capable at small parameter scales. At the same time, conventional cloud-centric deployment introduces challenges around data privacy, latency, and cost that are acute in operational technology and defence environments. Advances in model distillation, quantisation, and affordable edge accelerators now make local LLM inference on single-board computers feasible, but the high dimensionality of the configuration space makes identifying optimal deployments difficult without structured evaluation. Existing LLM-specific edge benchmarking efforts rely on CPU-only inference, poor coverage of genuine single-board computers, and generic evaluation tasks that lack multi-dimensional assessment of hardware effectiveness. This paper proposes a multi-dimensional benchmarking methodology that jointly evaluates inference performance and hardware efficiency across four IoT-suitable edge platform configurations testing single-board computers with the latest available hardware accelerators. Our results reveal the benefits of using hardware accelerators such as NPUs and GPUs, along with multi-dimensional evaluations quantifying the trade-offs between power efficiency, physical device size and token throughput; offering practical guidance for deploying generative AI in privacy-sensitive and connectivity-limited environments such as unmanned vehicles and portable, ruggedised operations.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24785
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
Renney, Harri
Trad, Fouad
Mattarock, Michael
Wood, Zena
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Performance
Large language models (LLMs) are becoming increasingly capable at small parameter scales. At the same time, conventional cloud-centric deployment introduces challenges around data privacy, latency, and cost that are acute in operational technology and defence environments. Advances in model distillation, quantisation, and affordable edge accelerators now make local LLM inference on single-board computers feasible, but the high dimensionality of the configuration space makes identifying optimal deployments difficult without structured evaluation. Existing LLM-specific edge benchmarking efforts rely on CPU-only inference, poor coverage of genuine single-board computers, and generic evaluation tasks that lack multi-dimensional assessment of hardware effectiveness. This paper proposes a multi-dimensional benchmarking methodology that jointly evaluates inference performance and hardware efficiency across four IoT-suitable edge platform configurations testing single-board computers with the latest available hardware accelerators. Our results reveal the benefits of using hardware accelerators such as NPUs and GPUs, along with multi-dimensional evaluations quantifying the trade-offs between power efficiency, physical device size and token throughput; offering practical guidance for deploying generative AI in privacy-sensitive and connectivity-limited environments such as unmanned vehicles and portable, ruggedised operations.
title Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2604.24785