Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACs
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Qizhe, Liang, Huawen, Gui, Yuchen, Zeng, Zhichen, He, Zerong, Tao, Linfeng, Wang, Xiaotian, Zhao, Letian, Zeng, Zhaoxi, Yuan, Wei, Wu, Wei, Jin, Xi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
di: Wu, Qizhe, et al.
Pubblicazione: (2024)
di: Wu, Qizhe, et al.
Pubblicazione: (2024)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
di: Wu, Qizhe, et al.
Pubblicazione: (2024)
di: Wu, Qizhe, et al.
Pubblicazione: (2024)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
di: Wang, Jiayi, et al.
Pubblicazione: (2026)
di: Wang, Jiayi, et al.
Pubblicazione: (2026)
FPPS: An FPGA-Based Point Cloud Processing System
di: Zhou, Xiaofeng, et al.
Pubblicazione: (2026)
di: Zhou, Xiaofeng, et al.
Pubblicazione: (2026)
Tensor Memory Engine: On-the-fly Data Reorganization for Ideal Locality
di: Hoornaert, Denis, et al.
Pubblicazione: (2026)
di: Hoornaert, Denis, et al.
Pubblicazione: (2026)
SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
di: Cai, Jingwei, et al.
Pubblicazione: (2025)
di: Cai, Jingwei, et al.
Pubblicazione: (2025)
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
di: Li, Jindong, et al.
Pubblicazione: (2024)
di: Li, Jindong, et al.
Pubblicazione: (2024)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
di: Wu, Yuting, et al.
Pubblicazione: (2023)
di: Wu, Yuting, et al.
Pubblicazione: (2023)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
di: Li, Tenglong, et al.
Pubblicazione: (2025)
di: Li, Tenglong, et al.
Pubblicazione: (2025)
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
di: Zheng, Keran, et al.
Pubblicazione: (2025)
di: Zheng, Keran, et al.
Pubblicazione: (2025)
Exploring the Versal AI Engine for 3D Gaussian Splatting
di: Shimamura, Kotaro, et al.
Pubblicazione: (2025)
di: Shimamura, Kotaro, et al.
Pubblicazione: (2025)
No One-Size-Fits-All: A Workload-Driven Characterization of Bit-Parallel vs. Bit-Serial Data Layouts for Processing-using-Memory
di: Zhang, Jingyao, et al.
Pubblicazione: (2025)
di: Zhang, Jingyao, et al.
Pubblicazione: (2025)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
Big-PERCIVAL: Exploring the Native Use of 64-Bit Posit Arithmetic in Scientific Computing
di: Mallasén, David, et al.
Pubblicazione: (2023)
di: Mallasén, David, et al.
Pubblicazione: (2023)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
di: Shin, Yongwon, et al.
Pubblicazione: (2024)
di: Shin, Yongwon, et al.
Pubblicazione: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
di: Sun, Xiaotian, et al.
Pubblicazione: (2024)
di: Sun, Xiaotian, et al.
Pubblicazione: (2024)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
di: Wang, Huizheng, et al.
Pubblicazione: (2025)
di: Wang, Huizheng, et al.
Pubblicazione: (2025)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
di: Li, Wanqian, et al.
Pubblicazione: (2024)
di: Li, Wanqian, et al.
Pubblicazione: (2024)
Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
di: Zhou, Weiyu, et al.
Pubblicazione: (2025)
di: Zhou, Weiyu, et al.
Pubblicazione: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
di: Shan, Haoxuan, et al.
Pubblicazione: (2025)
di: Shan, Haoxuan, et al.
Pubblicazione: (2025)
Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Capacity
di: Xue, Zi Yu, et al.
Pubblicazione: (2023)
di: Xue, Zi Yu, et al.
Pubblicazione: (2023)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
di: Wei, Jianyu, et al.
Pubblicazione: (2025)
di: Wei, Jianyu, et al.
Pubblicazione: (2025)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
di: Zhao, Wei, et al.
Pubblicazione: (2024)
di: Zhao, Wei, et al.
Pubblicazione: (2024)
A Bit Level Weight Reordering Strategy Based on Column Similarity to Explore Weight Sparsity in RRAM-based NN Accelerator
di: Yang, Weiping, et al.
Pubblicazione: (2025)
di: Yang, Weiping, et al.
Pubblicazione: (2025)
PyPIM: Integrating Digital Processing-in-Memory from Microarchitectural Design to Python Tensors
di: Leitersdorf, Orian, et al.
Pubblicazione: (2023)
di: Leitersdorf, Orian, et al.
Pubblicazione: (2023)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
di: Mo, Zhiwen, et al.
Pubblicazione: (2024)
di: Mo, Zhiwen, et al.
Pubblicazione: (2024)
A 64-Spin All-to-All CMOS Ising Machine with Landscape Perturbation Achieving 2.28 nJ/Edge-Bit Energy-to-Solution
di: Salim, Ahmet Yusuf, et al.
Pubblicazione: (2026)
di: Salim, Ahmet Yusuf, et al.
Pubblicazione: (2026)
Flexible Bit-Truncation Memory for Approximate Applications on the Edge
di: Oswald, William, et al.
Pubblicazione: (2025)
di: Oswald, William, et al.
Pubblicazione: (2025)
Shift-Left Techniques in Electronic Design Automation: A Survey
di: Wu, Xinyue, et al.
Pubblicazione: (2025)
di: Wu, Xinyue, et al.
Pubblicazione: (2025)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
di: Taka, Endri, et al.
Pubblicazione: (2025)
di: Taka, Endri, et al.
Pubblicazione: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
di: Ye, Hanchen, et al.
Pubblicazione: (2025)
di: Ye, Hanchen, et al.
Pubblicazione: (2025)
IMAGine: An In-Memory Accelerated GEMV Engine Overlay
di: Kabir, MD Arafat, et al.
Pubblicazione: (2024)
di: Kabir, MD Arafat, et al.
Pubblicazione: (2024)
Accelerating CRONet on AMD Versal AIE-ML Engines
di: Mhatre, Kaustubh, et al.
Pubblicazione: (2026)
di: Mhatre, Kaustubh, et al.
Pubblicazione: (2026)
Accelerating Sparse Graph Neural Networks with Tensor Core Optimization
di: Wu, Ka Wai
Pubblicazione: (2024)
di: Wu, Ka Wai
Pubblicazione: (2024)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
di: Liu, Xingyu, et al.
Pubblicazione: (2025)
di: Liu, Xingyu, et al.
Pubblicazione: (2025)
Optimizing and Exploring System Performance in Compact Processing-in-Memory-based Chips
di: Chen, Peilin, et al.
Pubblicazione: (2025)
di: Chen, Peilin, et al.
Pubblicazione: (2025)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
di: Gilbert, Michael, et al.
Pubblicazione: (2024)
di: Gilbert, Michael, et al.
Pubblicazione: (2024)
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
di: Wang, Huizheng, et al.
Pubblicazione: (2025)
di: Wang, Huizheng, et al.
Pubblicazione: (2025)
STAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar
di: Zhai, Yifeng, et al.
Pubblicazione: (2024)
di: Zhai, Yifeng, et al.
Pubblicazione: (2024)
A Novel Cost-Effective MIMO Architecture with Ray Antenna Array for Enhanced Wireless Communication Performance
di: Dong, Zhenjun, et al.
Pubblicazione: (2025)
di: Dong, Zhenjun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
di: Wu, Qizhe, et al.
Pubblicazione: (2024) -
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
di: Wu, Qizhe, et al.
Pubblicazione: (2024) -
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
di: Wang, Jiayi, et al.
Pubblicazione: (2026) -
FPPS: An FPGA-Based Point Cloud Processing System
di: Zhou, Xiaofeng, et al.
Pubblicazione: (2026) -
Tensor Memory Engine: On-the-fly Data Reorganization for Ideal Locality
di: Hoornaert, Denis, et al.
Pubblicazione: (2026)