Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | de Lima, João Paulo Cardoso, Dietrich, Marc, Castrillon, Jeronimo, Khan, Asif Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Full-Stack Optimization for CAM-Only DNN Inference
by: de Lima, João Paulo C., et al.
Published: (2024)
by: de Lima, João Paulo C., et al.
Published: (2024)
The Landscape of Compute-near-memory and Compute-in-memory: A Research and Commercial Overview
by: Khan, Asif Ali, et al.
Published: (2024)
by: Khan, Asif Ali, et al.
Published: (2024)
Count2Multiply: Reliable In-Memory High-Radix Counting
by: de Lima, João Paulo Cardoso, et al.
Published: (2024)
by: de Lima, João Paulo Cardoso, et al.
Published: (2024)
All-in-Memory Stochastic Computing using ReRAM
by: de Lima, João Paulo C., et al.
Published: (2025)
by: de Lima, João Paulo C., et al.
Published: (2025)
CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory Paradigms
by: Khan, Asif Ali, et al.
Published: (2022)
by: Khan, Asif Ali, et al.
Published: (2022)
Leveraging Stochastic Depth Training for Adaptive Inference
by: Korol, Guilherme, et al.
Published: (2025)
by: Korol, Guilherme, et al.
Published: (2025)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
by: Wang, Chenyu, et al.
Published: (2023)
by: Wang, Chenyu, et al.
Published: (2023)
Accelerating Sparse Graph Neural Networks with Tensor Core Optimization
by: Wu, Ka Wai
Published: (2024)
by: Wu, Ka Wai
Published: (2024)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
by: Su, Yuchao, et al.
Published: (2025)
by: Su, Yuchao, et al.
Published: (2025)
iEEG Seizure Detection with a Sparse Hyperdimensional Computing Accelerator
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
by: Kulp, Gabriel, et al.
Published: (2024)
by: Kulp, Gabriel, et al.
Published: (2024)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
by: Yang, Xiaoxuan, et al.
Published: (2025)
by: Yang, Xiaoxuan, et al.
Published: (2025)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
by: Wolters, Christopher, et al.
Published: (2024)
by: Wolters, Christopher, et al.
Published: (2024)
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
by: Morisaki, Keita
Published: (2026)
by: Morisaki, Keita
Published: (2026)
ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
by: Liu, Hongxiang, et al.
Published: (2025)
by: Liu, Hongxiang, et al.
Published: (2025)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
U-SWIM: Universal Selective Write-Verify for Computing-in-Memory Neural Accelerators
by: Yan, Zheyu, et al.
Published: (2023)
by: Yan, Zheyu, et al.
Published: (2023)
Accelerating Computer Architecture Simulation through Machine Learning
by: Ali, Wajid, et al.
Published: (2024)
by: Ali, Wajid, et al.
Published: (2024)
Designing Efficient LLM Accelerators for Edge Devices
by: Haris, Jude, et al.
Published: (2024)
by: Haris, Jude, et al.
Published: (2024)
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
by: Kim, Wonung, et al.
Published: (2025)
by: Kim, Wonung, et al.
Published: (2025)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
by: Chen, Yanru, et al.
Published: (2025)
by: Chen, Yanru, et al.
Published: (2025)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
by: Gupta, Neelesh, et al.
Published: (2026)
by: Gupta, Neelesh, et al.
Published: (2026)
Modeling and Simulating Emerging Memory Technologies: A Tutorial
by: Chen, Yun-Chih, et al.
Published: (2025)
by: Chen, Yun-Chih, et al.
Published: (2025)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
by: Shin, Duckgyu, et al.
Published: (2026)
by: Shin, Duckgyu, et al.
Published: (2026)
Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation
by: Park, Junyoung, et al.
Published: (2024)
by: Park, Junyoung, et al.
Published: (2024)
When Small Variations Become Big Failures: Reliability Challenges in Compute-in-Memory Neural Accelerators
by: Qin, Yifan, et al.
Published: (2026)
by: Qin, Yifan, et al.
Published: (2026)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
by: Cong, Rongqing, et al.
Published: (2024)
by: Cong, Rongqing, et al.
Published: (2024)
Mugi: Value Level Parallelism For Efficient LLMs
by: Price, Daniel, et al.
Published: (2026)
by: Price, Daniel, et al.
Published: (2026)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
by: İslamoğlu, Gamze, et al.
Published: (2023)
by: İslamoğlu, Gamze, et al.
Published: (2023)
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
by: Bhattacharya, Swastik, et al.
Published: (2025)
by: Bhattacharya, Swastik, et al.
Published: (2025)
FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
by: Hsieh, Fen-Yu, et al.
Published: (2025)
by: Hsieh, Fen-Yu, et al.
Published: (2025)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
by: Fang, Jiaxun, et al.
Published: (2025)
by: Fang, Jiaxun, et al.
Published: (2025)
An Efficient Data Reuse with Tile-Based Adaptive Stationary for Transformer Accelerators
by: Li, Tseng-Jen, et al.
Published: (2025)
by: Li, Tseng-Jen, et al.
Published: (2025)
SiHGNN: Leveraging Properties of Semantic Graphs for Efficient HGNN Acceleration
by: Xue, Runzhen, et al.
Published: (2024)
by: Xue, Runzhen, et al.
Published: (2024)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024)
by: Shao, Haikuo, et al.
Published: (2024)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs
by: Ahmadilivani, Mohammad Hasan, et al.
Published: (2026)
by: Ahmadilivani, Mohammad Hasan, et al.
Published: (2026)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
by: Wang, Miaoxin, et al.
Published: (2024)
by: Wang, Miaoxin, et al.
Published: (2024)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
Similar Items
-
Full-Stack Optimization for CAM-Only DNN Inference
by: de Lima, João Paulo C., et al.
Published: (2024) -
The Landscape of Compute-near-memory and Compute-in-memory: A Research and Commercial Overview
by: Khan, Asif Ali, et al.
Published: (2024) -
Count2Multiply: Reliable In-Memory High-Radix Counting
by: de Lima, João Paulo Cardoso, et al.
Published: (2024) -
All-in-Memory Stochastic Computing using ReRAM
by: de Lima, João Paulo C., et al.
Published: (2025) -
CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory Paradigms
by: Khan, Asif Ali, et al.
Published: (2022)