Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Chun-Ting, Mun, HanGyeol, Meng, Jian, Abdelfattah, Mohamed S., Seo, Jae-sun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BBS: Bi-directional Bit-level Sparsity for Deep Learning Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
TrIM, Triangular Input Movement Systolic Array for Convolutional Neural Networks: Dataflow and Analytical Modelling
von: Sestito, Cristian, et al.
Veröffentlicht: (2024)
von: Sestito, Cristian, et al.
Veröffentlicht: (2024)
Torch2Chip: An End-to-end Customizable Deep Neural Network Compression and Deployment Toolkit for Prototype Hardware Accelerator Design
von: Meng, Jian, et al.
Veröffentlicht: (2024)
von: Meng, Jian, et al.
Veröffentlicht: (2024)
Error Checking for Sparse Systolic Tensor Arrays
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
von: Cherezova, Natalia, et al.
Veröffentlicht: (2025)
von: Cherezova, Natalia, et al.
Veröffentlicht: (2025)
KAN-SAs: Efficient Acceleration of Kolmogorov-Arnold Networks on Systolic Arrays
von: Errabii, Sohaib, et al.
Veröffentlicht: (2025)
von: Errabii, Sohaib, et al.
Veröffentlicht: (2025)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
Stream-HLS: Towards Automatic Dataflow Acceleration
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
Exploration of Activation Fault Reliability in Quantized Systolic Array-Based DNN Accelerators
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)
Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks
von: Raja, Tejas
Veröffentlicht: (2024)
von: Raja, Tejas
Veröffentlicht: (2024)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
von: Ye, Hanchen, et al.
Veröffentlicht: (2025)
von: Ye, Hanchen, et al.
Veröffentlicht: (2025)
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
von: Ko, Sho, et al.
Veröffentlicht: (2024)
von: Ko, Sho, et al.
Veröffentlicht: (2024)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2024)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
Fast Cross-Operator Optimization of Attention Dataflow
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
Dataflow Optimized Reconfigurable Acceleration for FEM-based CFD Simulations
von: Kapetanakis, Anastassis, et al.
Veröffentlicht: (2024)
von: Kapetanakis, Anastassis, et al.
Veröffentlicht: (2024)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
Single-Event Upset Analysis of a Systolic Array based Deep Neural Network Accelerator
von: Jonckers, Naïn, et al.
Veröffentlicht: (2024)
von: Jonckers, Naïn, et al.
Veröffentlicht: (2024)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
von: Gilbert, Michael, et al.
Veröffentlicht: (2024)
von: Gilbert, Michael, et al.
Veröffentlicht: (2024)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
ONE-SA: Enabling Nonlinear Operations in Systolic Arrays for Efficient and Flexible Neural Network Inference
von: Sun, Ruiqi, et al.
Veröffentlicht: (2024)
von: Sun, Ruiqi, et al.
Veröffentlicht: (2024)
HF-NTT: Hazard-Free Dataflow Accelerator for Number Theoretic Transform
von: Meng, Xiangchen, et al.
Veröffentlicht: (2024)
von: Meng, Xiangchen, et al.
Veröffentlicht: (2024)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
Revet: A Language and Compiler for Dataflow Threads
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
von: Dai, Xilai, et al.
Veröffentlicht: (2024)
von: Dai, Xilai, et al.
Veröffentlicht: (2024)
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling
von: Jin, Jeongmin, et al.
Veröffentlicht: (2026)
von: Jin, Jeongmin, et al.
Veröffentlicht: (2026)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BBS: Bi-directional Bit-level Sparsity for Deep Learning Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024) -
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
von: Han, Meng, et al.
Veröffentlicht: (2023) -
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
von: Su, Yuchao, et al.
Veröffentlicht: (2025) -
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
von: Chen, Yuzong, et al.
Veröffentlicht: (2025) -
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)