CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration
Fuente:
arXiv
Saved in:
| Main Authors: | Ahn, Bas, Tao, Xingjian, Gomony, Manil Dev, Geilen, Marc, Corporaal, Henk |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs
by: Alexandris, Georgios, et al.
Published: (2025)
by: Alexandris, Georgios, et al.
Published: (2025)
LinkBo: An Adaptive Single-Wire, Low-Latency, and Fault-Tolerant Communications Interface for Variable-Distance Chip-to-Chip Systems
by: Ye, Bochen, et al.
Published: (2025)
by: Ye, Bochen, et al.
Published: (2025)
CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-based CIM Architectures
by: Qi, Yingjie, et al.
Published: (2025)
by: Qi, Yingjie, et al.
Published: (2025)
CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge Device
by: and, Yan-Cheng Guo, et al.
Published: (2025)
by: and, Yan-Cheng Guo, et al.
Published: (2025)
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
by: Chen, Jinwu, et al.
Published: (2026)
by: Chen, Jinwu, et al.
Published: (2026)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
by: Xue, Chenhao, et al.
Published: (2026)
by: Xue, Chenhao, et al.
Published: (2026)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
RCW-CIM: A Digital CIM-based LLM Accelerator with Read-Compute/Write
by: Guo, Yan-Cheng, et al.
Published: (2026)
by: Guo, Yan-Cheng, et al.
Published: (2026)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
by: Khabbazan, Bahareh, et al.
Published: (2025)
by: Khabbazan, Bahareh, et al.
Published: (2025)
A Review of SRAM-based Compute-in-Memory Circuits
by: Yoshioka, Kentaro, et al.
Published: (2024)
by: Yoshioka, Kentaro, et al.
Published: (2024)
LOREN: Low Rank-Based Code-Rate Adaptation in Neural Receivers
by: Van Bolderik, Bram, et al.
Published: (2026)
by: Van Bolderik, Bram, et al.
Published: (2026)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
by: Guo, Zeyu
Published: (2025)
by: Guo, Zeyu
Published: (2025)
GEM3D CIM General Purpose Matrix Computation Using 3D Integrated SRAM eDRAM Hybrid Compute In Memory on Memory Architecture
by: Chakraborty, Subhradip, et al.
Published: (2026)
by: Chakraborty, Subhradip, et al.
Published: (2026)
TL-nvSRAM-CIM: Ultra-High-Density Three-Level ReRAM-Assisted Computing-in-nvSRAM with DC-Power Free Restore and Ternary MAC Operations
by: Wang, Dengfeng, et al.
Published: (2023)
by: Wang, Dengfeng, et al.
Published: (2023)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
by: Andronic, Marta, et al.
Published: (2023)
by: Andronic, Marta, et al.
Published: (2023)
SRAM-PG: Power Delivery Network Benchmarks from SRAM Circuits
by: Shen, Shan, et al.
Published: (2024)
by: Shen, Shan, et al.
Published: (2024)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
by: Zhao, Shixin, et al.
Published: (2025)
by: Zhao, Shixin, et al.
Published: (2025)
SAIL: SRAM-Accelerated LLM Inference System with Lookup-Table-based GEMV
by: Zhang, Jingyao, et al.
Published: (2025)
by: Zhang, Jingyao, et al.
Published: (2025)
A 28nm 1.80Mb/mm2 Digital/Analog Hybrid SRAM-CIM Macro Using 2D-Weighted Capacitor Array for Complex Number Mac Operations
by: Konno, Shota, et al.
Published: (2025)
by: Konno, Shota, et al.
Published: (2025)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
by: Han, Wontak, et al.
Published: (2024)
by: Han, Wontak, et al.
Published: (2024)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
by: Lou, Binglei, et al.
Published: (2024)
by: Lou, Binglei, et al.
Published: (2024)
Voxel-CIM: An Efficient Compute-in-Memory Accelerator for Voxel-based Point Cloud Neural Networks
by: Lin, Xipeng, et al.
Published: (2024)
by: Lin, Xipeng, et al.
Published: (2024)
Acore-CIM: build accurate and reliable mixed-signal CIM cores with RISC-V controlled self-calibration
by: Numan, Omar, et al.
Published: (2025)
by: Numan, Omar, et al.
Published: (2025)
Energy-efficient SNN Architecture using 3nm FinFET Multiport SRAM-based CIM with Online Learning
by: Huijbregts, Lucas, et al.
Published: (2024)
by: Huijbregts, Lucas, et al.
Published: (2024)
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
by: Qin, Shantian, et al.
Published: (2025)
by: Qin, Shantian, et al.
Published: (2025)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
by: Xu, Weikai, et al.
Published: (2026)
by: Xu, Weikai, et al.
Published: (2026)
3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge Rendering
by: Huang, Wei-Hsing, et al.
Published: (2025)
by: Huang, Wei-Hsing, et al.
Published: (2025)
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
by: Lou, Binglei, et al.
Published: (2026)
by: Lou, Binglei, et al.
Published: (2026)
SparseLUT: Sparse Connectivity Optimization for Lookup Table-based Deep Neural Networks
by: Lou, Binglei, et al.
Published: (2025)
by: Lou, Binglei, et al.
Published: (2025)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
by: He, Zifan, et al.
Published: (2025)
by: He, Zifan, et al.
Published: (2025)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
by: Hong, Junguk, et al.
Published: (2026)
by: Hong, Junguk, et al.
Published: (2026)
DSLUT: An Asymmetric LUT and its Automatic Design Flow Based on Practical Functions
by: Yang, Moucheng, et al.
Published: (2025)
by: Yang, Moucheng, et al.
Published: (2025)
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
by: Pun, Junius, et al.
Published: (2025)
by: Pun, Junius, et al.
Published: (2025)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
by: Sonnino, Lorenzo, et al.
Published: (2023)
by: Sonnino, Lorenzo, et al.
Published: (2023)
CryptoSRAM: Enabling High-Throughput Cryptography on MCUs via In-SRAM Computing
by: Zhang, Jingyao, et al.
Published: (2025)
by: Zhang, Jingyao, et al.
Published: (2025)
Ultra8T: A Sub-Threshold 8T SRAM with Leakage Detection
by: Shen, Shan, et al.
Published: (2023)
by: Shen, Shan, et al.
Published: (2023)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
by: Huang, Zhirui, et al.
Published: (2025)
by: Huang, Zhirui, et al.
Published: (2025)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
by: Xuan, Zihao, et al.
Published: (2026)
by: Xuan, Zihao, et al.
Published: (2026)
Unicorn-CIM: Uncovering the Vulnerability and Improving the Resilience of High-Precision Compute-in-Memory
by: Li, Qiufeng, et al.
Published: (2025)
by: Li, Qiufeng, et al.
Published: (2025)
High-Level Surface Code Decoding via Parallel FFNNs on CIM Platforms
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Similar Items
-
A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs
by: Alexandris, Georgios, et al.
Published: (2025) -
LinkBo: An Adaptive Single-Wire, Low-Latency, and Fault-Tolerant Communications Interface for Variable-Distance Chip-to-Chip Systems
by: Ye, Bochen, et al.
Published: (2025) -
CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-based CIM Architectures
by: Qi, Yingjie, et al.
Published: (2025) -
CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge Device
by: and, Yan-Cheng Guo, et al.
Published: (2025) -
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
by: Chen, Jinwu, et al.
Published: (2026)