Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Carrigg, Kieran, de Vries, Sigur, Sadough, Amirhossein, van Gerven, Marcel
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909041121296384
author Carrigg, Kieran
de Vries, Sigur
Sadough, Amirhossein
van Gerven, Marcel
author_facet Carrigg, Kieran
de Vries, Sigur
Sadough, Amirhossein
van Gerven, Marcel
contents Vision Transformers (ViTs) achieve state-of-the-art performance on challenging vision tasks, but their deployment on edge devices is severely hindered by the computational complexity and global reduction bottleneck imposed by layer normalization. Recent methods attempt to bypass this by replacing normalization layers with hardware-friendly scalar approximations. However, these homogeneous replacements do not optimally fit to all layers' behaviour and rely on expensive model retraining. In this work, we propose a highly efficient, hardware-aware framework that utilizes genetic programming (GP) to evolve heterogeneous, layer-specific scalar functions directly from pre-trained weights. Coupled with a novel post-training re-alignment strategy, our approach eliminates the need to retrain models from scratch entirely. Our evolved expressions accurately approximate the target normalization behaviours, capturing $91.6\%$ of the variance ($R^2$) compared to only $70.2\%$ for homogeneous baselines, allowing our modified architecture to recover $84.25\%$ Top-1 ImageNet-1K accuracy in only 20 epochs. By preserving this performance while eliminating the global reduction bottleneck, our approach establishes a highly favourable trade-off between arithmetic complexity and off-chip memory traffic, removing a primary barrier to the efficient deployment of ViTs on edge accelerators.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14047
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation
Carrigg, Kieran
de Vries, Sigur
Sadough, Amirhossein
van Gerven, Marcel
Computer Vision and Pattern Recognition
Hardware Architecture
Vision Transformers (ViTs) achieve state-of-the-art performance on challenging vision tasks, but their deployment on edge devices is severely hindered by the computational complexity and global reduction bottleneck imposed by layer normalization. Recent methods attempt to bypass this by replacing normalization layers with hardware-friendly scalar approximations. However, these homogeneous replacements do not optimally fit to all layers' behaviour and rely on expensive model retraining. In this work, we propose a highly efficient, hardware-aware framework that utilizes genetic programming (GP) to evolve heterogeneous, layer-specific scalar functions directly from pre-trained weights. Coupled with a novel post-training re-alignment strategy, our approach eliminates the need to retrain models from scratch entirely. Our evolved expressions accurately approximate the target normalization behaviours, capturing $91.6\%$ of the variance ($R^2$) compared to only $70.2\%$ for homogeneous baselines, allowing our modified architecture to recover $84.25\%$ Top-1 ImageNet-1K accuracy in only 20 epochs. By preserving this performance while eliminating the global reduction bottleneck, our approach establishes a highly favourable trade-off between arithmetic complexity and off-chip memory traffic, removing a primary barrier to the efficient deployment of ViTs on edge accelerators.
title Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation
topic Computer Vision and Pattern Recognition
Hardware Architecture
url https://arxiv.org/abs/2605.14047