The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908949302738944 |
|---|---|
| author | Geens, Robin De Schouwer, Jonas Verhelst, Marian Tambe, Thierry |
| author_facet | Geens, Robin De Schouwer, Jonas Verhelst, Marian Tambe, Thierry |
| contents | The Hardware Lottery posits that research directions are dictated by available silicon compute platforms. We identify a derivative phenomenon, the Hyperscale Lottery, where model architectures are optimized for cloud throughput at the expense of algorithmic efficiency. While State-Space Models (SSMs) such as Mamba were lauded for their linear complexity, ideal for edge intelligence, their evolution from Mamba-1 to Mamba-3 reveals a systematic divergence from edge-native efficiency. We demonstrate that Mamba-3's architectural changes, designed to saturate hyperscale GPUs, impose a significant edge penalty: a 28% latency increase at 880M parameters, worsening to 48% for 15M-parameter models. We argue for decoupling cloud-scale saturation strategies from core architectural design to preserve the viability of single-user, real-time edge intelligence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_07935 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency Geens, Robin De Schouwer, Jonas Verhelst, Marian Tambe, Thierry Hardware Architecture The Hardware Lottery posits that research directions are dictated by available silicon compute platforms. We identify a derivative phenomenon, the Hyperscale Lottery, where model architectures are optimized for cloud throughput at the expense of algorithmic efficiency. While State-Space Models (SSMs) such as Mamba were lauded for their linear complexity, ideal for edge intelligence, their evolution from Mamba-1 to Mamba-3 reveals a systematic divergence from edge-native efficiency. We demonstrate that Mamba-3's architectural changes, designed to saturate hyperscale GPUs, impose a significant edge penalty: a 28% latency increase at 880M parameters, worsening to 48% for 15M-parameter models. We argue for decoupling cloud-scale saturation strategies from core architectural design to preserve the viability of single-user, real-time edge intelligence. |
| title | The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency |
| topic | Hardware Architecture |
| url | https://arxiv.org/abs/2604.07935 |