Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Wei, Chiyue, Duan, Bowen, Guo, Cong, Zhang, Jingyang, Song, Qingyue, Li, Hai "Helen", Chen, Yiran |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
por: Wei, Chiyue, et al.
Publicado: (2025)
por: Wei, Chiyue, et al.
Publicado: (2025)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
por: Guo, Cong, et al.
Publicado: (2025)
por: Guo, Cong, et al.
Publicado: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
por: Shan, Haoxuan, et al.
Publicado: (2025)
por: Shan, Haoxuan, et al.
Publicado: (2025)
FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
por: Fu, Yuzhe, et al.
Publicado: (2025)
por: Fu, Yuzhe, et al.
Publicado: (2025)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
por: Cheng, Feng, et al.
Publicado: (2025)
por: Cheng, Feng, et al.
Publicado: (2025)
Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
por: Wei, Chiyue, et al.
Publicado: (2025)
por: Wei, Chiyue, et al.
Publicado: (2025)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
por: Duan, Bowen, et al.
Publicado: (2026)
por: Duan, Bowen, et al.
Publicado: (2026)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
por: Li, Tenglong, et al.
Publicado: (2024)
por: Li, Tenglong, et al.
Publicado: (2024)
Sparsity-Aware Hardware-Software Co-Design of Spiking Neural Networks: An Overview
por: Aliyev, Ilkin, et al.
Publicado: (2024)
por: Aliyev, Ilkin, et al.
Publicado: (2024)
CAMformer: Associative Memory is All You Need
por: Molom-Ochir, Tergel, et al.
Publicado: (2025)
por: Molom-Ochir, Tergel, et al.
Publicado: (2025)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
por: Li, Tenglong, et al.
Publicado: (2025)
por: Li, Tenglong, et al.
Publicado: (2025)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
por: Manoni, Simone, et al.
Publicado: (2025)
por: Manoni, Simone, et al.
Publicado: (2025)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
por: Wang, Huizheng, et al.
Publicado: (2025)
por: Wang, Huizheng, et al.
Publicado: (2025)
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
por: Wang, Yitu, et al.
Publicado: (2023)
por: Wang, Yitu, et al.
Publicado: (2023)
Leveraging Recurrent Patterns in Graph Accelerators
por: Rahimi, Masoud, et al.
Publicado: (2025)
por: Rahimi, Masoud, et al.
Publicado: (2025)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
por: Guo, Zeyu
Publicado: (2025)
por: Guo, Zeyu
Publicado: (2025)
Surrogates, Spikes, and Sparsity: Performance Analysis and Characterization of SNN Hyperparameters on Hardware
por: Aliyev, Ilkin, et al.
Publicado: (2026)
por: Aliyev, Ilkin, et al.
Publicado: (2026)
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
por: Li, Tenglong, et al.
Publicado: (2026)
por: Li, Tenglong, et al.
Publicado: (2026)
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
por: Molom-Ochir, Tergel, et al.
Publicado: (2024)
por: Molom-Ochir, Tergel, et al.
Publicado: (2024)
A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator with In-Situ Regulation for Energy-Efficient Spiking Neural Networks
por: Kao, Shih-Hang, et al.
Publicado: (2026)
por: Kao, Shih-Hang, et al.
Publicado: (2026)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
por: Duan, Cenlin, et al.
Publicado: (2024)
por: Duan, Cenlin, et al.
Publicado: (2024)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
por: Lu, Jinming, et al.
Publicado: (2025)
por: Lu, Jinming, et al.
Publicado: (2025)
Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
por: Duan, Cenlin, et al.
Publicado: (2025)
por: Duan, Cenlin, et al.
Publicado: (2025)
Xpikeformer: Hybrid Analog-Digital Hardware Acceleration for Spiking Transformers
por: Song, Zihang, et al.
Publicado: (2024)
por: Song, Zihang, et al.
Publicado: (2024)
An Event-Driven Spiking Compute-In-Memory Macro based on SOT-MRAM
por: Yu, Deyang, et al.
Publicado: (2025)
por: Yu, Deyang, et al.
Publicado: (2025)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
por: Fan, Zhenkun, et al.
Publicado: (2026)
por: Fan, Zhenkun, et al.
Publicado: (2026)
Look-Up Table based Neural Network Hardware
por: Sen, Ovishake, et al.
Publicado: (2024)
por: Sen, Ovishake, et al.
Publicado: (2024)
A Fully-Configurable Open-Source Software-Defined Digital Quantized Spiking Neural Core Architecture
por: Matinizadeh, Shadi, et al.
Publicado: (2024)
por: Matinizadeh, Shadi, et al.
Publicado: (2024)
UniSpike: Accelerating Spiking Neural Networks on Neuromorphic Systems via Eliminating Address Redundancy
por: Xing, Qinghui, et al.
Publicado: (2026)
por: Xing, Qinghui, et al.
Publicado: (2026)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
por: Yoon, Jieon, et al.
Publicado: (2026)
por: Yoon, Jieon, et al.
Publicado: (2026)
A Bit Level Weight Reordering Strategy Based on Column Similarity to Explore Weight Sparsity in RRAM-based NN Accelerator
por: Yang, Weiping, et al.
Publicado: (2025)
por: Yang, Weiping, et al.
Publicado: (2025)
Gaze into the Pattern: Characterizing Spatial Patterns with Internal Temporal Correlations for Hardware Prefetching
por: Chen, Zixiao, et al.
Publicado: (2024)
por: Chen, Zixiao, et al.
Publicado: (2024)
RealProbe: An Automated and Lightweight Performance Profiler for In-FPGA Execution of High-Level Synthesis Designs
por: Kim, Jiho, et al.
Publicado: (2025)
por: Kim, Jiho, et al.
Publicado: (2025)
Automated Physical Design Watermarking Leveraging Graph Neural Networks
por: Zhang, Ruisi, et al.
Publicado: (2024)
por: Zhang, Ruisi, et al.
Publicado: (2024)
A 28.6 mJ/iter Stable Diffusion Processor for Text-to-Image Generation with Patch Similarity-based Sparsity Augmentation and Text-based Mixed-Precision
por: Choi, Jiwon, et al.
Publicado: (2024)
por: Choi, Jiwon, et al.
Publicado: (2024)
PULSE: Parametric Hardware Units for Low-power Sparsity-Aware Convolution Engine
por: Aliyev, Ilkin, et al.
Publicado: (2024)
por: Aliyev, Ilkin, et al.
Publicado: (2024)
VIKIN: A Reconfigurable Accelerator for KANs and MLPs with Two-Stage Sparsity Support
por: Ou, Wenhui, et al.
Publicado: (2026)
por: Ou, Wenhui, et al.
Publicado: (2026)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
por: Dai, Xilai, et al.
Publicado: (2024)
por: Dai, Xilai, et al.
Publicado: (2024)
Shooting Neutrons at Neurons: Radiation Testing of a Spiking Neural Network on Flash-Based FPGAs
por: Nijsink, Wim, et al.
Publicado: (2026)
por: Nijsink, Wim, et al.
Publicado: (2026)
Ejemplares similares
-
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
por: Wei, Chiyue, et al.
Publicado: (2025) -
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
por: Guo, Cong, et al.
Publicado: (2025) -
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
por: Shan, Haoxuan, et al.
Publicado: (2025) -
FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
por: Fu, Yuzhe, et al.
Publicado: (2025) -
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
por: Cheng, Feng, et al.
Publicado: (2025)