Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Luoyang, Jiang, Jiwen, Ding, Yifeng, Li, Fengfa, Song, Yan, Zhang, Haifeng, Ying, Jian, Ren, Lei, Zhan, Kun, Chen, Wei, Xie, Yan, Deng, Cheng
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910018639495168
author Sun, Luoyang
Jiang, Jiwen
Ding, Yifeng
Li, Fengfa
Song, Yan
Zhang, Haifeng
Ying, Jian
Ren, Lei
Zhan, Kun
Chen, Wei
Xie, Yan
Deng, Cheng
author_facet Sun, Luoyang
Jiang, Jiwen
Ding, Yifeng
Li, Fengfa
Song, Yan
Zhang, Haifeng
Ying, Jian
Ren, Lei
Zhan, Kun
Chen, Wei
Xie, Yan
Deng, Cheng
contents Vision-Language-Action Models (VLAs) have emerged as a key paradigm of Physical AI and are increasingly deployed in autonomous vehicles, robots, and smart spaces. In these resource-constrained on-device settings, selecting an appropriate large language model (LLM) backbone is a critical challenge: models must balance accuracy with strict inference latency and hardware efficiency constraints. This makes hardware-software co-design a game-changing requirement for on-device LLM deployment, where each hardware platform demands a tailored architectural solution. We propose a hardware co-design law that jointly captures model accuracy and inference performance. Specifically, we model training loss as an explicit function of architectural hyperparameters and characterise inference latency via roofline modelling. We empirically evaluate 1,942 candidate architectures on NVIDIA Jetson Orin, training 170 selected models for 10B tokens each to fit a scaling law relating architecture to training loss. By coupling this scaling law with latency modelling, we establish a direct accuracy-latency correspondence and identify the Pareto frontier for hardware co-designed LLMs. We further formulate architecture search as a joint optimisation over precision and performance, deriving feasible design regions under industrial hardware and application budgets. Our approach reduces architecture selection from months to days. At the same latency as Qwen2.5-0.5B on the target hardware, our co-designed architecture achieves 19.42% lower perplexity on WikiText-2. To our knowledge, this is the first principled and operational framework for hardware co-design scaling laws in on-device LLM deployment. We will make the code and related checkpoints publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10377
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
Sun, Luoyang
Jiang, Jiwen
Ding, Yifeng
Li, Fengfa
Song, Yan
Zhang, Haifeng
Ying, Jian
Ren, Lei
Zhan, Kun
Chen, Wei
Xie, Yan
Deng, Cheng
Machine Learning
Computation and Language
I.2.7
Vision-Language-Action Models (VLAs) have emerged as a key paradigm of Physical AI and are increasingly deployed in autonomous vehicles, robots, and smart spaces. In these resource-constrained on-device settings, selecting an appropriate large language model (LLM) backbone is a critical challenge: models must balance accuracy with strict inference latency and hardware efficiency constraints. This makes hardware-software co-design a game-changing requirement for on-device LLM deployment, where each hardware platform demands a tailored architectural solution. We propose a hardware co-design law that jointly captures model accuracy and inference performance. Specifically, we model training loss as an explicit function of architectural hyperparameters and characterise inference latency via roofline modelling. We empirically evaluate 1,942 candidate architectures on NVIDIA Jetson Orin, training 170 selected models for 10B tokens each to fit a scaling law relating architecture to training loss. By coupling this scaling law with latency modelling, we establish a direct accuracy-latency correspondence and identify the Pareto frontier for hardware co-designed LLMs. We further formulate architecture search as a joint optimisation over precision and performance, deriving feasible design regions under industrial hardware and application budgets. Our approach reduces architecture selection from months to days. At the same latency as Qwen2.5-0.5B on the target hardware, our co-designed architecture achieves 19.42% lower perplexity on WikiText-2. To our knowledge, this is the first principled and operational framework for hardware co-design scaling laws in on-device LLM deployment. We will make the code and related checkpoints publicly available.
title Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
topic Machine Learning
Computation and Language
I.2.7
url https://arxiv.org/abs/2602.10377