The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geens, Robin, De Schouwer, Jonas, Verhelst, Marian, Tambe, Thierry
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908949302738944
author Geens, Robin
De Schouwer, Jonas
Verhelst, Marian
Tambe, Thierry
author_facet Geens, Robin
De Schouwer, Jonas
Verhelst, Marian
Tambe, Thierry
contents The Hardware Lottery posits that research directions are dictated by available silicon compute platforms. We identify a derivative phenomenon, the Hyperscale Lottery, where model architectures are optimized for cloud throughput at the expense of algorithmic efficiency. While State-Space Models (SSMs) such as Mamba were lauded for their linear complexity, ideal for edge intelligence, their evolution from Mamba-1 to Mamba-3 reveals a systematic divergence from edge-native efficiency. We demonstrate that Mamba-3's architectural changes, designed to saturate hyperscale GPUs, impose a significant edge penalty: a 28% latency increase at 880M parameters, worsening to 48% for 15M-parameter models. We argue for decoupling cloud-scale saturation strategies from core architectural design to preserve the viability of single-user, real-time edge intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07935
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
Geens, Robin
De Schouwer, Jonas
Verhelst, Marian
Tambe, Thierry
Hardware Architecture
The Hardware Lottery posits that research directions are dictated by available silicon compute platforms. We identify a derivative phenomenon, the Hyperscale Lottery, where model architectures are optimized for cloud throughput at the expense of algorithmic efficiency. While State-Space Models (SSMs) such as Mamba were lauded for their linear complexity, ideal for edge intelligence, their evolution from Mamba-1 to Mamba-3 reveals a systematic divergence from edge-native efficiency. We demonstrate that Mamba-3's architectural changes, designed to saturate hyperscale GPUs, impose a significant edge penalty: a 28% latency increase at 880M parameters, worsening to 48% for 15M-parameter models. We argue for decoupling cloud-scale saturation strategies from core architectural design to preserve the viability of single-user, real-time edge intelligence.
title The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
topic Hardware Architecture
url https://arxiv.org/abs/2604.07935