Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mesa, Alejandro Ruiz y, Korol, Guilherme, Riesterer, Moritz, de Lima, João Paulo Cardoso, Castrillon, Jeronimo
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914317936361472
author Mesa, Alejandro Ruiz y
Korol, Guilherme
Riesterer, Moritz
de Lima, João Paulo Cardoso
Castrillon, Jeronimo
author_facet Mesa, Alejandro Ruiz y
Korol, Guilherme
Riesterer, Moritz
de Lima, João Paulo Cardoso
Castrillon, Jeronimo
contents LLM deployment on resource-constrained edge devices faces severe latency constraints, particularly in real-time applications where delayed responses can compromise safety or usability. Among many approaches to mitigate the inefficiencies of sequential token-by-token generation, Speculative Decoding (SD) has emerged as a promising technique. However, SD at the edge is hindered by two major challenges: (1) integrating SD into a compiler-based workflow without sacrificing performance or programmability, and (2) exploiting the heterogeneous compute resources of modern SoCs through carefully designed partitioning strategies. This work addresses these challenges by using an analytical cost model that explores heterogeneous hardware configurations and guides coarse-grained partitioning of LLM subgraphs, particularly with edge-typical short input sequence lengths. The cost model predicts when speculative sampling and heterogeneous execution are jointly beneficial and is validated on an edge device featuring a hexacore Cortex-A CPU and a Mali GPU, revealing up to 1.68$\times$ speedup for translation tasks, closely matching analytic expectations.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08060
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices
Mesa, Alejandro Ruiz y
Korol, Guilherme
Riesterer, Moritz
de Lima, João Paulo Cardoso
Castrillon, Jeronimo
Machine Learning
LLM deployment on resource-constrained edge devices faces severe latency constraints, particularly in real-time applications where delayed responses can compromise safety or usability. Among many approaches to mitigate the inefficiencies of sequential token-by-token generation, Speculative Decoding (SD) has emerged as a promising technique. However, SD at the edge is hindered by two major challenges: (1) integrating SD into a compiler-based workflow without sacrificing performance or programmability, and (2) exploiting the heterogeneous compute resources of modern SoCs through carefully designed partitioning strategies. This work addresses these challenges by using an analytical cost model that explores heterogeneous hardware configurations and guides coarse-grained partitioning of LLM subgraphs, particularly with edge-typical short input sequence lengths. The cost model predicts when speculative sampling and heterogeneous execution are jointly beneficial and is validated on an edge device featuring a hexacore Cortex-A CPU and a Mali GPU, revealing up to 1.68$\times$ speedup for translation tasks, closely matching analytic expectations.
title Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices
topic Machine Learning
url https://arxiv.org/abs/2602.08060