Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
Fuente:
arXiv
Saved in:
| Main Authors: | Geens, Robin, Symons, Arne, Verhelst, Marian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
by: Fang, Chao, et al.
Published: (2024)
by: Fang, Chao, et al.
Published: (2024)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
by: Symons, Arne, et al.
Published: (2022)
by: Symons, Arne, et al.
Published: (2022)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
by: Mei, Linyan, et al.
Published: (2022)
by: Mei, Linyan, et al.
Published: (2022)
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
by: Colleman, Steven, et al.
Published: (2024)
by: Colleman, Steven, et al.
Published: (2024)
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
by: Jung, Victor J. B., et al.
Published: (2023)
by: Jung, Victor J. B., et al.
Published: (2023)
Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators
by: Zniber, Alaa, et al.
Published: (2025)
by: Zniber, Alaa, et al.
Published: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)
by: Dumoulin, Joren, et al.
Published: (2025)
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
by: Yi, Xiaoling, et al.
Published: (2025)
by: Yi, Xiaoling, et al.
Published: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
by: Verhelst, Marian, et al.
Published: (2025)
by: Verhelst, Marian, et al.
Published: (2025)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
by: Yi, Xiaoling, et al.
Published: (2026)
by: Yi, Xiaoling, et al.
Published: (2026)
iEEG Seizure Detection with a Sparse Hyperdimensional Computing Accelerator
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
by: Houshmand, Pouya, et al.
Published: (2024)
by: Houshmand, Pouya, et al.
Published: (2024)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
by: Antonio, Ryan Albert, et al.
Published: (2025)
by: Antonio, Ryan Albert, et al.
Published: (2025)
Decoupled Control Flow and Data Access in RISC-V GPGPUs
by: Sarda, Giuseppe M., et al.
Published: (2025)
by: Sarda, Giuseppe M., et al.
Published: (2025)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
by: Sarda, Giuseppe M., et al.
Published: (2024)
by: Sarda, Giuseppe M., et al.
Published: (2024)
Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models
by: Odemuyiwa, Toluwanimi O., et al.
Published: (2026)
by: Odemuyiwa, Toluwanimi O., et al.
Published: (2026)
MC$^2$A: Enabling Algorithm-Hardware Co-Design for Efficient Markov Chain Monte Carlo Acceleration
by: Zhao, Shirui, et al.
Published: (2025)
by: Zhao, Shirui, et al.
Published: (2025)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025)
by: Chen, Yuzong, et al.
Published: (2025)
Analog or Digital In-memory Computing? Benchmarking through Quantitative Modeling
by: Sun, Jiacong, et al.
Published: (2024)
by: Sun, Jiacong, et al.
Published: (2024)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
by: Kong, Fanchen, et al.
Published: (2025)
by: Kong, Fanchen, et al.
Published: (2025)
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
by: Langarita, Rubén, et al.
Published: (2025)
by: Langarita, Rubén, et al.
Published: (2025)
OpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling
by: Yi, Xiaoling, et al.
Published: (2024)
by: Yi, Xiaoling, et al.
Published: (2024)
FLICKER: A Fine-Grained Contribution-Aware Accelerator for Real-Time 3D Gaussian Splatting
by: Ou, Wenhui, et al.
Published: (2026)
by: Ou, Wenhui, et al.
Published: (2026)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
by: Deng, Yunhao, et al.
Published: (2025)
by: Deng, Yunhao, et al.
Published: (2025)
DreamRAM: A Fine-Grained Configurable Design Space Modeling Tool for Custom 3D Die-Stacked DRAM
by: Cai, Victor, et al.
Published: (2025)
by: Cai, Victor, et al.
Published: (2025)
Sectored DRAM: A Practical Energy-Efficient and High-Performance Fine-Grained DRAM Architecture
by: Olgun, Ataberk, et al.
Published: (2022)
by: Olgun, Ataberk, et al.
Published: (2022)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
by: Xuan, Zihao, et al.
Published: (2026)
by: Xuan, Zihao, et al.
Published: (2026)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
by: Yoon, Hyunsung, et al.
Published: (2026)
by: Yoon, Hyunsung, et al.
Published: (2026)
Modeling Analog-Digital-Converter Energy and Area for Compute-In-Memory Accelerator Design
by: Andrulis, Tanner, et al.
Published: (2024)
by: Andrulis, Tanner, et al.
Published: (2024)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
by: Shi, Man, et al.
Published: (2024)
by: Shi, Man, et al.
Published: (2024)
FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
by: Jia, Shuao, et al.
Published: (2025)
by: Jia, Shuao, et al.
Published: (2025)
Accelerating Sensor Fusion in Neuromorphic Computing: A Case Study on Loihi-2
by: Isik, Murat, et al.
Published: (2024)
by: Isik, Murat, et al.
Published: (2024)
Piccolo: Large-Scale Graph Processing with Fine-Grained In-Memory Scatter-Gather
by: Shin, Changmin, et al.
Published: (2025)
by: Shin, Changmin, et al.
Published: (2025)
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
by: Nguyen, Hoa, et al.
Published: (2025)
by: Nguyen, Hoa, et al.
Published: (2025)
Similar Items
-
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025) -
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
by: Geens, Robin, et al.
Published: (2026) -
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026) -
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
by: Fang, Chao, et al.
Published: (2024) -
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
by: Symons, Arne, et al.
Published: (2022)