Saved in:
| Main Authors: | Wu, Qizhe, Liang, Huawen, Gui, Yuchen, Zeng, Zhichen, He, Zerong, Tao, Linfeng, Wang, Xiaotian, Zhao, Letian, Zeng, Zhaoxi, Yuan, Wei, Wu, Wei, Jin, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.06342 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
by: Wu, Qizhe, et al.
Published: (2024)
by: Wu, Qizhe, et al.
Published: (2024)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
by: Wu, Qizhe, et al.
Published: (2024)
by: Wu, Qizhe, et al.
Published: (2024)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
FPPS: An FPGA-Based Point Cloud Processing System
by: Zhou, Xiaofeng, et al.
Published: (2026)
by: Zhou, Xiaofeng, et al.
Published: (2026)
Tensor Memory Engine: On-the-fly Data Reorganization for Ideal Locality
by: Hoornaert, Denis, et al.
Published: (2026)
by: Hoornaert, Denis, et al.
Published: (2026)
SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
by: Cai, Jingwei, et al.
Published: (2025)
by: Cai, Jingwei, et al.
Published: (2025)
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
by: Li, Jindong, et al.
Published: (2024)
by: Li, Jindong, et al.
Published: (2024)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
by: Li, Tenglong, et al.
Published: (2025)
by: Li, Tenglong, et al.
Published: (2025)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
by: Wu, Yuting, et al.
Published: (2023)
by: Wu, Yuting, et al.
Published: (2023)
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
by: Zheng, Keran, et al.
Published: (2025)
by: Zheng, Keran, et al.
Published: (2025)
Exploring the Versal AI Engine for 3D Gaussian Splatting
by: Shimamura, Kotaro, et al.
Published: (2025)
by: Shimamura, Kotaro, et al.
Published: (2025)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
No One-Size-Fits-All: A Workload-Driven Characterization of Bit-Parallel vs. Bit-Serial Data Layouts for Processing-using-Memory
by: Zhang, Jingyao, et al.
Published: (2025)
by: Zhang, Jingyao, et al.
Published: (2025)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
by: Sun, Xiaotian, et al.
Published: (2024)
by: Sun, Xiaotian, et al.
Published: (2024)
Big-PERCIVAL: Exploring the Native Use of 64-Bit Posit Arithmetic in Scientific Computing
by: Mallasén, David, et al.
Published: (2023)
by: Mallasén, David, et al.
Published: (2023)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
by: Li, Wanqian, et al.
Published: (2024)
by: Li, Wanqian, et al.
Published: (2024)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
by: Shin, Yongwon, et al.
Published: (2024)
by: Shin, Yongwon, et al.
Published: (2024)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
by: Shan, Haoxuan, et al.
Published: (2025)
by: Shan, Haoxuan, et al.
Published: (2025)
Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
by: Zhou, Weiyu, et al.
Published: (2025)
by: Zhou, Weiyu, et al.
Published: (2025)
Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Capacity
by: Xue, Zi Yu, et al.
Published: (2023)
by: Xue, Zi Yu, et al.
Published: (2023)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
by: Wei, Jianyu, et al.
Published: (2025)
by: Wei, Jianyu, et al.
Published: (2025)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
A Bit Level Weight Reordering Strategy Based on Column Similarity to Explore Weight Sparsity in RRAM-based NN Accelerator
by: Yang, Weiping, et al.
Published: (2025)
by: Yang, Weiping, et al.
Published: (2025)
Shift-Left Techniques in Electronic Design Automation: A Survey
by: Wu, Xinyue, et al.
Published: (2025)
by: Wu, Xinyue, et al.
Published: (2025)
PyPIM: Integrating Digital Processing-in-Memory from Microarchitectural Design to Python Tensors
by: Leitersdorf, Orian, et al.
Published: (2023)
by: Leitersdorf, Orian, et al.
Published: (2023)
A 64-Spin All-to-All CMOS Ising Machine with Landscape Perturbation Achieving 2.28 nJ/Edge-Bit Energy-to-Solution
by: Salim, Ahmet Yusuf, et al.
Published: (2026)
by: Salim, Ahmet Yusuf, et al.
Published: (2026)
Fast and Practical Strassen's Matrix Multiplication using FPGAs
by: Ahmad, Afzal, et al.
Published: (2024)
by: Ahmad, Afzal, et al.
Published: (2024)
Allo: A Programming Model for Composable Accelerator Design
by: Chen, Hongzheng, et al.
Published: (2024)
by: Chen, Hongzheng, et al.
Published: (2024)
Flexible Bit-Truncation Memory for Approximate Applications on the Edge
by: Oswald, William, et al.
Published: (2025)
by: Oswald, William, et al.
Published: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
Accelerating Sparse Graph Neural Networks with Tensor Core Optimization
by: Wu, Ka Wai
Published: (2024)
by: Wu, Ka Wai
Published: (2024)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
by: Ye, Hanchen, et al.
Published: (2025)
by: Ye, Hanchen, et al.
Published: (2025)
IMAGine: An In-Memory Accelerated GEMV Engine Overlay
by: Kabir, MD Arafat, et al.
Published: (2024)
by: Kabir, MD Arafat, et al.
Published: (2024)
WISP: Image Segmentation-Based Whitespace Diagnosis for Optimal Rectilinear Floorplanning
by: Zhao, Xiaotian, et al.
Published: (2025)
by: Zhao, Xiaotian, et al.
Published: (2025)
Accelerating CRONet on AMD Versal AIE-ML Engines
by: Mhatre, Kaustubh, et al.
Published: (2026)
by: Mhatre, Kaustubh, et al.
Published: (2026)
TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
by: Shen, Chaoyao, et al.
Published: (2026)
by: Shen, Chaoyao, et al.
Published: (2026)
Similar Items
-
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
by: Wu, Qizhe, et al.
Published: (2024) -
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
by: Wu, Qizhe, et al.
Published: (2024) -
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
by: Wang, Jiayi, et al.
Published: (2026) -
FPPS: An FPGA-Based Point Cloud Processing System
by: Zhou, Xiaofeng, et al.
Published: (2026) -
Tensor Memory Engine: On-the-fly Data Reorganization for Ideal Locality
by: Hoornaert, Denis, et al.
Published: (2026)