FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Hsieh, Fen-Yu, Teng, Yun-Chang, Hong, Ding-Yong, Wu, Jan-Jan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RNM-TD3: N:M Semi-structured Sparse Reinforcement Learning From Scratch
by: Vrce, Isam, et al.
Published: (2026)
by: Vrce, Isam, et al.
Published: (2026)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
by: Dong, Jiale, et al.
Published: (2025)
by: Dong, Jiale, et al.
Published: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
by: Wang, Peipei, et al.
Published: (2025)
by: Wang, Peipei, et al.
Published: (2025)
ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA
by: Lyu, Shengzhe, et al.
Published: (2026)
by: Lyu, Shengzhe, et al.
Published: (2026)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
by: Gao, Yizhao, et al.
Published: (2024)
by: Gao, Yizhao, et al.
Published: (2024)
Energy-Efficient FPGA Framework for Non-Quantized Convolutional Neural Networks
by: Athanasiadis, Angelos, et al.
Published: (2025)
by: Athanasiadis, Angelos, et al.
Published: (2025)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
by: Yu, Feng, et al.
Published: (2026)
by: Yu, Feng, et al.
Published: (2026)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
by: Wang, Miaoxin, et al.
Published: (2024)
by: Wang, Miaoxin, et al.
Published: (2024)
KLLM: Fast LLM Inference with K-Means Quantization
by: Wu, Xueying, et al.
Published: (2025)
by: Wu, Xueying, et al.
Published: (2025)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
by: Yin, Chenyang, et al.
Published: (2025)
by: Yin, Chenyang, et al.
Published: (2025)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
by: Shin, Duckgyu, et al.
Published: (2026)
by: Shin, Duckgyu, et al.
Published: (2026)
A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE
by: Okubo, Ikumi, et al.
Published: (2024)
by: Okubo, Ikumi, et al.
Published: (2024)
STI-SNN: A 0.14 GOPS/W/PE Single-Timestep Inference FPGA-based SNN Accelerator with Algorithm and Hardware Co-Design
by: Wang, Kainan, et al.
Published: (2025)
by: Wang, Kainan, et al.
Published: (2025)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
by: Andronic, Marta, et al.
Published: (2023)
by: Andronic, Marta, et al.
Published: (2023)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024)
by: Shao, Haikuo, et al.
Published: (2024)
Evaluating Four FPGA-accelerated Space Use Cases based on Neural Network Algorithms for On-board Inference
by: Antunes, Pedro, et al.
Published: (2026)
by: Antunes, Pedro, et al.
Published: (2026)
Efficient FPGA Implementation of Time-Domain Popcount for Low-Complexity Machine Learning
by: Duan, Shengyu, et al.
Published: (2025)
by: Duan, Shengyu, et al.
Published: (2025)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
by: Klhufek, Jan, et al.
Published: (2024)
by: Klhufek, Jan, et al.
Published: (2024)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
by: Oh, Hyunwoo, et al.
Published: (2026)
by: Oh, Hyunwoo, et al.
Published: (2026)
SparseDPD: A Sparse Neural Network-based Digital Predistortion FPGA Accelerator for RF Power Amplifier Linearization
by: Versluis, Manno, et al.
Published: (2025)
by: Versluis, Manno, et al.
Published: (2025)
T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
by: Oh, Hyunwoo, et al.
Published: (2025)
by: Oh, Hyunwoo, et al.
Published: (2025)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
by: Wu, Junyi, et al.
Published: (2025)
by: Wu, Junyi, et al.
Published: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
by: Zhang, Zehuan, et al.
Published: (2026)
by: Zhang, Zehuan, et al.
Published: (2026)
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
by: Xie, Yanyue, et al.
Published: (2024)
by: Xie, Yanyue, et al.
Published: (2024)
Design and Implementation of BNN-Based Object Detection on FPGA
by: Zhao, Xuyu, et al.
Published: (2026)
by: Zhao, Xuyu, et al.
Published: (2026)
Exploring FPGA designs for MX and beyond
by: Samson, Ebby, et al.
Published: (2024)
by: Samson, Ebby, et al.
Published: (2024)
Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models
by: Vahdatpour, Mohammad Saleh, et al.
Published: (2026)
by: Vahdatpour, Mohammad Saleh, et al.
Published: (2026)
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
by: Fyon, Arthur, et al.
Published: (2026)
by: Fyon, Arthur, et al.
Published: (2026)
Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
by: Habermann, Tobias, et al.
Published: (2026)
by: Habermann, Tobias, et al.
Published: (2026)
InstantFT: An FPGA-Based Runtime Subsecond Fine-tuning of CNN Models
by: Sugiura, Keisuke, et al.
Published: (2025)
by: Sugiura, Keisuke, et al.
Published: (2025)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
by: Parthasarathy, Rishab
Published: (2025)
by: Parthasarathy, Rishab
Published: (2025)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
by: Ding, Hong, et al.
Published: (2025)
by: Ding, Hong, et al.
Published: (2025)
Similar Items
-
RNM-TD3: N:M Semi-structured Sparse Reinforcement Learning From Scratch
by: Vrce, Isam, et al.
Published: (2026) -
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024) -
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
by: Dong, Jiale, et al.
Published: (2025) -
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025) -
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)