Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
Fuente:
arXiv
Saved in:
| Main Authors: | Renney, Harri, Trad, Fouad, Mattarock, Michael, Wood, Zena |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024)
by: Qararyah, Fareed, et al.
Published: (2024)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
by: Chakraborty, Abhinaba, et al.
Published: (2025)
by: Chakraborty, Abhinaba, et al.
Published: (2025)
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
by: Hu, Ziyu, et al.
Published: (2025)
by: Hu, Ziyu, et al.
Published: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI Accelerators
by: Tang, Xinsheng, et al.
Published: (2026)
by: Tang, Xinsheng, et al.
Published: (2026)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
by: Lei, Jianlong, et al.
Published: (2026)
by: Lei, Jianlong, et al.
Published: (2026)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
by: Panigrahy, Deepak, et al.
Published: (2026)
by: Panigrahy, Deepak, et al.
Published: (2026)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
by: Ibeid, Huda, et al.
Published: (2025)
by: Ibeid, Huda, et al.
Published: (2025)
UPMEM Unleashed: Software Secrets for Speed
by: Chmielewski, Krystian, et al.
Published: (2025)
by: Chmielewski, Krystian, et al.
Published: (2025)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025)
by: Luo, Weile, et al.
Published: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
by: Aqasizade, Hossein, et al.
Published: (2024)
by: Aqasizade, Hossein, et al.
Published: (2024)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
by: Szafarczyk, Robert, et al.
Published: (2025)
by: Szafarczyk, Robert, et al.
Published: (2025)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Parallelizing a modern GPU simulator
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
by: Wang, Chengyue, et al.
Published: (2025)
by: Wang, Chengyue, et al.
Published: (2025)
Exploiting long vectors with a CFD code: a co-design show case
by: Blancafort, Marc, et al.
Published: (2024)
by: Blancafort, Marc, et al.
Published: (2024)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
by: Wadhwa, Eashan, et al.
Published: (2024)
by: Wadhwa, Eashan, et al.
Published: (2024)
Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry
by: Chen, Ziji, et al.
Published: (2025)
by: Chen, Ziji, et al.
Published: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
by: Kumar, Deepak, et al.
Published: (2025)
by: Kumar, Deepak, et al.
Published: (2025)
COMET: Neural Cost Model Explanation Framework
by: Chaudhary, Isha, et al.
Published: (2023)
by: Chaudhary, Isha, et al.
Published: (2023)
Evaluating the Potential of In-Memory Processing to Accelerate Homomorphic Encryption
by: Mwaisela, Mpoki, et al.
Published: (2024)
by: Mwaisela, Mpoki, et al.
Published: (2024)
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
by: John, Chelsea Maria, et al.
Published: (2024)
by: John, Chelsea Maria, et al.
Published: (2024)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
by: Adamopoulos, Dionysios, et al.
Published: (2025)
by: Adamopoulos, Dionysios, et al.
Published: (2025)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
Accelerator-as-a-Service in Public Clouds: An Intra-Host Traffic Management View for Performance Isolation in the Wild
by: Zhao, Jiechen, et al.
Published: (2024)
by: Zhao, Jiechen, et al.
Published: (2024)
Vector-Centric Machine Learning Systems: A Cross-Stack Approach
by: Jiang, Wenqi
Published: (2025)
by: Jiang, Wenqi
Published: (2025)
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
by: Shi, Tianyao, et al.
Published: (2025)
by: Shi, Tianyao, et al.
Published: (2025)
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
by: Zhou, Keren, et al.
Published: (2025)
by: Zhou, Keren, et al.
Published: (2025)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
by: Feng, Weigang, et al.
Published: (2025)
by: Feng, Weigang, et al.
Published: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
by: Xu, Weihong, et al.
Published: (2025)
by: Xu, Weihong, et al.
Published: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
by: Kwak, Hyunseok, et al.
Published: (2025)
by: Kwak, Hyunseok, et al.
Published: (2025)
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
by: Grailoo, M., et al.
Published: (2026)
by: Grailoo, M., et al.
Published: (2026)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
by: Kubwimana, Benjamin, et al.
Published: (2025)
by: Kubwimana, Benjamin, et al.
Published: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
by: Qiu, Tong Dong, et al.
Published: (2023)
by: Qiu, Tong Dong, et al.
Published: (2023)
Similar Items
-
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025) -
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026) -
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024) -
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
by: Chakraborty, Abhinaba, et al.
Published: (2025) -
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
by: Hu, Ziyu, et al.
Published: (2025)