VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Helal, Shereef, Garcia-Ortiz, Alberto, Bamberg, Lennart |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SAFFIRA: a Framework for Assessing the Reliability of Systolic-Array-Based DNN Accelerators
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
Explainable AI-Guided Efficient Approximate DNN Generation for Multi-Pod Systolic Arrays
von: Siddique, Ayesha, et al.
Veröffentlicht: (2025)
von: Siddique, Ayesha, et al.
Veröffentlicht: (2025)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
von: Müller, Mika Markus, et al.
Veröffentlicht: (2025)
von: Müller, Mika Markus, et al.
Veröffentlicht: (2025)
eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
von: Bamberg, Lennart, et al.
Veröffentlicht: (2025)
von: Bamberg, Lennart, et al.
Veröffentlicht: (2025)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
von: Jeong, Geonhwa, et al.
Veröffentlicht: (2024)
von: Jeong, Geonhwa, et al.
Veröffentlicht: (2024)
EXION: Exploiting Inter- and Intra-Iteration Output Sparsity for Diffusion Models
von: Heo, Jaehoon, et al.
Veröffentlicht: (2025)
von: Heo, Jaehoon, et al.
Veröffentlicht: (2025)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
von: Mueller, Lion, et al.
Veröffentlicht: (2025)
von: Mueller, Lion, et al.
Veröffentlicht: (2025)
KAN-SAs: Efficient Acceleration of Kolmogorov-Arnold Networks on Systolic Arrays
von: Errabii, Sohaib, et al.
Veröffentlicht: (2025)
von: Errabii, Sohaib, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
von: Ma, Shaobo, et al.
Veröffentlicht: (2025)
von: Ma, Shaobo, et al.
Veröffentlicht: (2025)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
von: Chen, Chun-Ting, et al.
Veröffentlicht: (2025)
von: Chen, Chun-Ting, et al.
Veröffentlicht: (2025)
Exploration of Activation Fault Reliability in Quantized Systolic Array-Based DNN Accelerators
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
The Role of Advanced Computer Architectures in Accelerating Artificial Intelligence Workloads
von: Amin, Shahid, et al.
Veröffentlicht: (2025)
von: Amin, Shahid, et al.
Veröffentlicht: (2025)
SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling
von: Jin, Jeongmin, et al.
Veröffentlicht: (2026)
von: Jin, Jeongmin, et al.
Veröffentlicht: (2026)
Periodic Online Testing for Sparse Systolic Tensor Arrays
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2025)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2025)
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
von: Meng, Chang, et al.
Veröffentlicht: (2026)
von: Meng, Chang, et al.
Veröffentlicht: (2026)
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2024)
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2024)
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture
von: Prakash, Shvetank, et al.
Veröffentlicht: (2025)
von: Prakash, Shvetank, et al.
Veröffentlicht: (2025)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
von: Cherezova, Natalia, et al.
Veröffentlicht: (2025)
von: Cherezova, Natalia, et al.
Veröffentlicht: (2025)
ALADIN: Accuracy-Latency-Aware Design-space Inference Analysis for Embedded AI Accelerators
von: Baldi, T., et al.
Veröffentlicht: (2026)
von: Baldi, T., et al.
Veröffentlicht: (2026)
ElasticAI: Creating and Deploying Energy-Efficient Deep Learning Accelerator for Pervasive Computing
von: Qian, Chao, et al.
Veröffentlicht: (2024)
von: Qian, Chao, et al.
Veröffentlicht: (2024)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
von: Nair, Harideep, et al.
Veröffentlicht: (2024)
von: Nair, Harideep, et al.
Veröffentlicht: (2024)
TrIM, Triangular Input Movement Systolic Array for Convolutional Neural Networks: Dataflow and Analytical Modelling
von: Sestito, Cristian, et al.
Veröffentlicht: (2024)
von: Sestito, Cristian, et al.
Veröffentlicht: (2024)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
Heterogeneous Acceleration Pipeline for Recommendation System Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2022)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2022)
Dynamic Sparse Attention: Access Patterns and Architecture
von: Levy, Noam
Veröffentlicht: (2026)
von: Levy, Noam
Veröffentlicht: (2026)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
von: Jung, Alexander Louis-Ferdinand, et al.
Veröffentlicht: (2024)
von: Jung, Alexander Louis-Ferdinand, et al.
Veröffentlicht: (2024)
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
von: Lin, Yujun, et al.
Veröffentlicht: (2025)
von: Lin, Yujun, et al.
Veröffentlicht: (2025)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
von: Demirkiran, Cansu, et al.
Veröffentlicht: (2023)
von: Demirkiran, Cansu, et al.
Veröffentlicht: (2023)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
Differentiable Initialization-Accelerated CPU-GPU Hybrid Combinatorial Scheduling
von: Liu, Mingju, et al.
Veröffentlicht: (2026)
von: Liu, Mingju, et al.
Veröffentlicht: (2026)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
von: Kabir, Ehsan, et al.
Veröffentlicht: (2024)
von: Kabir, Ehsan, et al.
Veröffentlicht: (2024)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
AutoHLS: Learning to Accelerate Design Space Exploration for HLS Designs
von: Ahmed, Md Rubel, et al.
Veröffentlicht: (2024)
von: Ahmed, Md Rubel, et al.
Veröffentlicht: (2024)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
AdAM: Adaptive Fault-Tolerant Approximate Multiplier for Edge DNN Accelerators
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SAFFIRA: a Framework for Assessing the Reliability of Systolic-Array-Based DNN Accelerators
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024) -
Explainable AI-Guided Efficient Approximate DNN Generation for Multi-Pod Systolic Arrays
von: Siddique, Ayesha, et al.
Veröffentlicht: (2025) -
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
von: Müller, Mika Markus, et al.
Veröffentlicht: (2025) -
eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
von: Bamberg, Lennart, et al.
Veröffentlicht: (2025) -
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
von: Jeong, Geonhwa, et al.
Veröffentlicht: (2024)