The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
Fuente:
arXiv
Saved in:
| Main Authors: | Geens, Robin, De Schouwer, Jonas, Verhelst, Marian, Tambe, Thierry |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025)
by: Chen, Yuzong, et al.
Published: (2025)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
by: Fang, Chao, et al.
Published: (2024)
by: Fang, Chao, et al.
Published: (2024)
How to keep pushing ML accelerator performance? Know your rooflines!
by: Verhelst, Marian, et al.
Published: (2025)
by: Verhelst, Marian, et al.
Published: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)
by: Dumoulin, Joren, et al.
Published: (2025)
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
by: Symons, Arne, et al.
Published: (2022)
by: Symons, Arne, et al.
Published: (2022)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
by: Houshmand, Pouya, et al.
Published: (2024)
by: Houshmand, Pouya, et al.
Published: (2024)
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
by: Colleman, Steven, et al.
Published: (2024)
by: Colleman, Steven, et al.
Published: (2024)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
by: Mei, Linyan, et al.
Published: (2022)
by: Mei, Linyan, et al.
Published: (2022)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
by: Sarda, Giuseppe M., et al.
Published: (2024)
by: Sarda, Giuseppe M., et al.
Published: (2024)
Decoupled Control Flow and Data Access in RISC-V GPGPUs
by: Sarda, Giuseppe M., et al.
Published: (2025)
by: Sarda, Giuseppe M., et al.
Published: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
by: Yi, Xiaoling, et al.
Published: (2025)
by: Yi, Xiaoling, et al.
Published: (2025)
iEEG Seizure Detection with a Sparse Hyperdimensional Computing Accelerator
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
by: Yi, Xiaoling, et al.
Published: (2026)
by: Yi, Xiaoling, et al.
Published: (2026)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
Analog or Digital In-memory Computing? Benchmarking through Quantitative Modeling
by: Sun, Jiacong, et al.
Published: (2024)
by: Sun, Jiacong, et al.
Published: (2024)
Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators
by: Zniber, Alaa, et al.
Published: (2025)
by: Zniber, Alaa, et al.
Published: (2025)
The Future of Memory: Limits and Opportunities
by: Dayo, Samuel, et al.
Published: (2025)
by: Dayo, Samuel, et al.
Published: (2025)
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
by: Jung, Victor J. B., et al.
Published: (2023)
by: Jung, Victor J. B., et al.
Published: (2023)
LLM-FSM: Scaling Large Language Models for Finite-State Reasoning in RTL Code Generation
by: Wu, Yuheng, et al.
Published: (2026)
by: Wu, Yuheng, et al.
Published: (2026)
Lottery BP: Unlocking Quantum Error Decoding at Scale
by: Zhu, Yanzhang, et al.
Published: (2026)
by: Zhu, Yanzhang, et al.
Published: (2026)
Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models
by: Odemuyiwa, Toluwanimi O., et al.
Published: (2026)
by: Odemuyiwa, Toluwanimi O., et al.
Published: (2026)
OpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling
by: Yi, Xiaoling, et al.
Published: (2024)
by: Yi, Xiaoling, et al.
Published: (2024)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
by: Deng, Yunhao, et al.
Published: (2025)
by: Deng, Yunhao, et al.
Published: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
by: Kong, Fanchen, et al.
Published: (2025)
by: Kong, Fanchen, et al.
Published: (2025)
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
by: Antonio, Ryan Albert, et al.
Published: (2025)
by: Antonio, Ryan Albert, et al.
Published: (2025)
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
by: Ko, Sho, et al.
Published: (2025)
by: Ko, Sho, et al.
Published: (2025)
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
by: Huang, Mingqiang, et al.
Published: (2024)
by: Huang, Mingqiang, et al.
Published: (2024)
MC$^2$A: Enabling Algorithm-Hardware Co-Design for Efficient Markov Chain Monte Carlo Acceleration
by: Zhao, Shirui, et al.
Published: (2025)
by: Zhao, Shirui, et al.
Published: (2025)
Computing-In-Memory Aware Model Adaption For Edge Devices
by: Lin, Ming-Han, et al.
Published: (2025)
by: Lin, Ming-Han, et al.
Published: (2025)
OpenGCRAM: An Open-Source Gain Cell Compiler Enabling Design-Space Exploration for AI Workloads
by: Wang, Xinxin, et al.
Published: (2025)
by: Wang, Xinxin, et al.
Published: (2025)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
by: Shi, Man, et al.
Published: (2024)
by: Shi, Man, et al.
Published: (2024)
How to Increase Energy Efficiency with a Single Linux Command
by: Jelvani, Alborz, et al.
Published: (2025)
by: Jelvani, Alborz, et al.
Published: (2025)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
by: Majeed, Ashiyana Abdul, et al.
Published: (2025)
by: Majeed, Ashiyana Abdul, et al.
Published: (2025)
Towards Memory Specialization: A Case for Long-Term and Short-Term RAM
by: Li, Peijing, et al.
Published: (2025)
by: Li, Peijing, et al.
Published: (2025)
EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
by: Bai, Kangbo, et al.
Published: (2025)
by: Bai, Kangbo, et al.
Published: (2025)
Modeling PFAS in Semiconductor Manufacturing to Quantify Trade-offs in Energy Efficiency and Environmental Impact of Computing Systems
by: Elgamal, Mariam, et al.
Published: (2025)
by: Elgamal, Mariam, et al.
Published: (2025)
Similar Items
-
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
by: Geens, Robin, et al.
Published: (2025) -
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025) -
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026) -
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025) -
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
by: Fang, Chao, et al.
Published: (2024)