A Low-Power Sparse Deep Learning Accelerator with Optimized Data Reuse
Fuente:
arXiv
Guardado en:
| Autores principales: | Hsu, Kai-Chieh, Chang, Tian-Sheuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Efficient Data Reuse with Tile-Based Adaptive Stationary for Transformer Accelerators
por: Li, Tseng-Jen, et al.
Publicado: (2025)
por: Li, Tseng-Jen, et al.
Publicado: (2025)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
por: Hsiung, Ching-Lin, et al.
Publicado: (2025)
por: Hsiung, Ching-Lin, et al.
Publicado: (2025)
VESTA: A Versatile SNN-Based Transformer Accelerator with Unified PEs for Multiple Computational Layers
por: Chen, Ching-Yao, et al.
Publicado: (2025)
por: Chen, Ching-Yao, et al.
Publicado: (2025)
VitaLLM: A Versatile, Ultra-Compact Ternary LLM Accelerator with Dependency-Aware Scheduling
por: Lin, Zi-Wei, et al.
Publicado: (2026)
por: Lin, Zi-Wei, et al.
Publicado: (2026)
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
por: Lin, Zi-Wei, et al.
Publicado: (2026)
por: Lin, Zi-Wei, et al.
Publicado: (2026)
Hardware Efficient Accelerator for Spiking Transformer With Reconfigurable Parallel Time Step Computing
por: Chen, Bo-Yu, et al.
Publicado: (2025)
por: Chen, Bo-Yu, et al.
Publicado: (2025)
A Low-Power Streaming Speech Enhancement Accelerator For Edge Devices
por: Wu, Ci-Hao, et al.
Publicado: (2025)
por: Wu, Ci-Hao, et al.
Publicado: (2025)
Dynamic Gradient Sparse Update for Edge Training
por: Li, I-Hsuan, et al.
Publicado: (2025)
por: Li, I-Hsuan, et al.
Publicado: (2025)
Enhancing Finite State Machine Design Automation with Large Language Models and Prompt Engineering Techniques
por: Lin, Qun-Kai, et al.
Publicado: (2025)
por: Lin, Qun-Kai, et al.
Publicado: (2025)
RCW-CIM: A Digital CIM-based LLM Accelerator with Read-Compute/Write
por: Guo, Yan-Cheng, et al.
Publicado: (2026)
por: Guo, Yan-Cheng, et al.
Publicado: (2026)
ESSR: An 8K@30FPS Super-Resolution Accelerator With Edge Selective Network
por: Hsu, Chih-Chia, et al.
Publicado: (2025)
por: Hsu, Chih-Chia, et al.
Publicado: (2025)
Computing-In-Memory Aware Model Adaption For Edge Devices
por: Lin, Ming-Han, et al.
Publicado: (2025)
por: Lin, Ming-Han, et al.
Publicado: (2025)
A 129FPS Full HD Real-Time Accelerator for 3D Gaussian Splatting
por: Chang, Fang-Chi, et al.
Publicado: (2026)
por: Chang, Fang-Chi, et al.
Publicado: (2026)
CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge Device
por: and, Yan-Cheng Guo, et al.
Publicado: (2025)
por: and, Yan-Cheng Guo, et al.
Publicado: (2025)
A 71.2-$μ$W Speech Recognition Accelerator with Recurrent Spiking Neural Network
por: Yang, Chih-Chyau, et al.
Publicado: (2025)
por: Yang, Chih-Chyau, et al.
Publicado: (2025)
A 14uJ/Decision Keyword Spotting Accelerator with In-SRAM-Computing and On Chip Learning for Customization
por: Chiang, Yu-Hsiang, et al.
Publicado: (2022)
por: Chiang, Yu-Hsiang, et al.
Publicado: (2022)
ASC: Adaptive Scale Feature Map Compression for Deep Neural Network
por: Yao, Yuan, et al.
Publicado: (2023)
por: Yao, Yuan, et al.
Publicado: (2023)
A Logic-Reuse Approach to Nibble-based Multiplier Design for Low Power Vector Computing
por: Chowdhury, Md Rownak Hossain, et al.
Publicado: (2026)
por: Chowdhury, Md Rownak Hossain, et al.
Publicado: (2026)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
por: Taka, Endri, et al.
Publicado: (2025)
por: Taka, Endri, et al.
Publicado: (2025)
ACNPU: A 4.75TOPS/W 1080P@30FPS Super Resolution Accelerator with Decoupled Asymmetric Convolution
por: Yang, Tun-Hao, et al.
Publicado: (2023)
por: Yang, Tun-Hao, et al.
Publicado: (2023)
GTA: a new General Tensor Accelerator with Better Area Efficiency and Data Reuse
por: Ai, Chenyang, et al.
Publicado: (2024)
por: Ai, Chenyang, et al.
Publicado: (2024)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
por: Guo, Cong, et al.
Publicado: (2025)
por: Guo, Cong, et al.
Publicado: (2025)
HyDRA: Deadline and Reuse-Aware Cacheability for Hardware Accelerators
por: Agarwal, Ayushi, et al.
Publicado: (2026)
por: Agarwal, Ayushi, et al.
Publicado: (2026)
SparseDPD: A Sparse Neural Network-based Digital Predistortion FPGA Accelerator for RF Power Amplifier Linearization
por: Versluis, Manno, et al.
Publicado: (2025)
por: Versluis, Manno, et al.
Publicado: (2025)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
por: Su, Yuchao, et al.
Publicado: (2025)
por: Su, Yuchao, et al.
Publicado: (2025)
A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator with In-Situ Regulation for Energy-Efficient Spiking Neural Networks
por: Kao, Shih-Hang, et al.
Publicado: (2026)
por: Kao, Shih-Hang, et al.
Publicado: (2026)
LogicSparse: Enabling Engine-Free Unstructured Sparsity for Quantised Deep-learning Accelerators
por: Li, Changhong, et al.
Publicado: (2025)
por: Li, Changhong, et al.
Publicado: (2025)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
por: Yi, Xiaoling, et al.
Publicado: (2026)
por: Yi, Xiaoling, et al.
Publicado: (2026)
Enhanced Hybrid Temporal Computing Using Deterministic Summations for Ultra-Low-Power Accelerators
por: Sachdeva, Sachin, et al.
Publicado: (2025)
por: Sachdeva, Sachin, et al.
Publicado: (2025)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
por: Wu, Junyi, et al.
Publicado: (2025)
por: Wu, Junyi, et al.
Publicado: (2025)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
por: Yoon, Jieon, et al.
Publicado: (2026)
por: Yoon, Jieon, et al.
Publicado: (2026)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
por: Xu, Haocheng, et al.
Publicado: (2024)
por: Xu, Haocheng, et al.
Publicado: (2024)
SigDLA: A Deep Learning Accelerator Extension for Signal Processing
por: Fu, Fangfa, et al.
Publicado: (2024)
por: Fu, Fangfa, et al.
Publicado: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
por: Li, Zhengke, et al.
Publicado: (2025)
por: Li, Zhengke, et al.
Publicado: (2025)
TeAAL: A Declarative Framework for Modeling Sparse Tensor Accelerators
por: Nayak, Nandeeka, et al.
Publicado: (2023)
por: Nayak, Nandeeka, et al.
Publicado: (2023)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
por: Tong, Jianming, et al.
Publicado: (2024)
por: Tong, Jianming, et al.
Publicado: (2024)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
por: Ferry, Corentin, et al.
Publicado: (2024)
por: Ferry, Corentin, et al.
Publicado: (2024)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
por: Yoon, Hyunsung, et al.
Publicado: (2026)
por: Yoon, Hyunsung, et al.
Publicado: (2026)
Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Capacity
por: Xue, Zi Yu, et al.
Publicado: (2023)
por: Xue, Zi Yu, et al.
Publicado: (2023)
Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators
por: Sakhuja, Chirag, et al.
Publicado: (2024)
por: Sakhuja, Chirag, et al.
Publicado: (2024)
Ejemplares similares
-
An Efficient Data Reuse with Tile-Based Adaptive Stationary for Transformer Accelerators
por: Li, Tseng-Jen, et al.
Publicado: (2025) -
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
por: Hsiung, Ching-Lin, et al.
Publicado: (2025) -
VESTA: A Versatile SNN-Based Transformer Accelerator with Unified PEs for Multiple Computational Layers
por: Chen, Ching-Yao, et al.
Publicado: (2025) -
VitaLLM: A Versatile, Ultra-Compact Ternary LLM Accelerator with Dependency-Aware Scheduling
por: Lin, Zi-Wei, et al.
Publicado: (2026) -
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
por: Lin, Zi-Wei, et al.
Publicado: (2026)