Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
Fuente:
arXiv
Saved in:
| Main Authors: | Klhufek, Jan, Safar, Miroslav, Mrazek, Vojtech, Vasicek, Zdenek, Sekanina, Lukas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AxMED: Formal Analysis and Automated Design of Approximate Median Filters using BDDs
by: Mrazek, Vojtech, et al.
Published: (2025)
by: Mrazek, Vojtech, et al.
Published: (2025)
TRAPTI: Time-Resolved Analysis for SRAM Banking and Power Gating Optimization in Embedded Transformer Inference
by: Klhufek, Jan, et al.
Published: (2026)
by: Klhufek, Jan, et al.
Published: (2026)
Late Breaking Result: FPGA-Based Emulation and Fault Injection for CNN Inference Accelerators
by: Masar, Filip, et al.
Published: (2025)
by: Masar, Filip, et al.
Published: (2025)
ObfAx: Obfuscation and IP Piracy Detection in Approximate Circuits
by: Sekanina, Lukas, et al.
Published: (2026)
by: Sekanina, Lukas, et al.
Published: (2026)
Evolutionary Approximation of Ternary Neurons for On-sensor Printed Neural Networks
by: Mrazek, Vojtech, et al.
Published: (2024)
by: Mrazek, Vojtech, et al.
Published: (2024)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
ApproxGNN: A Pretrained GNN for Parameter Prediction in Design Space Exploration for Approximate Computing
by: Vlcek, Ondrej, et al.
Published: (2025)
by: Vlcek, Ondrej, et al.
Published: (2025)
Torch2Chip: An End-to-end Customizable Deep Neural Network Compression and Deployment Toolkit for Prototype Hardware Accelerator Design
by: Meng, Jian, et al.
Published: (2024)
by: Meng, Jian, et al.
Published: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
by: Yu, Zhewen, et al.
Published: (2024)
by: Yu, Zhewen, et al.
Published: (2024)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Energy-Aware Deep Learning on Resource-Constrained Hardware
by: Millar, Josh, et al.
Published: (2025)
by: Millar, Josh, et al.
Published: (2025)
Hardware-Efficient Photonic Tensor Core: Accelerating Deep Neural Networks with Structured Compression
by: Ning, Shupeng, et al.
Published: (2025)
by: Ning, Shupeng, et al.
Published: (2025)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
by: Killian, Earl
Published: (2026)
by: Killian, Earl
Published: (2026)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
by: Hsiung, Ching-Lin, et al.
Published: (2025)
by: Hsiung, Ching-Lin, et al.
Published: (2025)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
by: Zhang, Zehuan, et al.
Published: (2024)
by: Zhang, Zehuan, et al.
Published: (2024)
Optical Computing for Deep Neural Network Acceleration: Foundations, Recent Developments, and Emerging Directions
by: Pasricha, Sudeep
Published: (2024)
by: Pasricha, Sudeep
Published: (2024)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
by: Mueller, Lion, et al.
Published: (2025)
by: Mueller, Lion, et al.
Published: (2025)
Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
Accelerating PoT Quantization on Edge Devices
by: Saha, Rappy, et al.
Published: (2024)
by: Saha, Rappy, et al.
Published: (2024)
Histogram-Equalized Quantization for logic-gated Residual Neural Networks
by: Nguyen, Van Thien, et al.
Published: (2025)
by: Nguyen, Van Thien, et al.
Published: (2025)
Exploring the Limitations of Kolmogorov-Arnold Networks in Classification: Insights to Software Training and Hardware Implementation
by: Tran, Van Duy, et al.
Published: (2024)
by: Tran, Van Duy, et al.
Published: (2024)
TreeLUT: An Efficient Alternative to Deep Neural Networks for Inference Acceleration Using Gradient Boosted Decision Trees
by: Khataei, Alireza, et al.
Published: (2025)
by: Khataei, Alireza, et al.
Published: (2025)
Hardware-Aware Fine-Tuning of Spiking Q-Networks on the SpiNNaker2 Neuromorphic Platform
by: Arfa, Sirine, et al.
Published: (2025)
by: Arfa, Sirine, et al.
Published: (2025)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
by: İslamoğlu, Gamze, et al.
Published: (2023)
by: İslamoğlu, Gamze, et al.
Published: (2023)
Accelerating Sparse Graph Neural Networks with Tensor Core Optimization
by: Wu, Ka Wai
Published: (2024)
by: Wu, Ka Wai
Published: (2024)
Sustainable Transformer Neural Network Acceleration with Stochastic Photonic Computing
by: Afifi, S., et al.
Published: (2026)
by: Afifi, S., et al.
Published: (2026)
MCEL: Margin-Based Cross-Entropy Loss for Error-Tolerant Quantized Neural Networks
by: Yayla, Mikail, et al.
Published: (2026)
by: Yayla, Mikail, et al.
Published: (2026)
Designing Approximate Arithmetic Circuits with Combined Error Constraints
by: Češka, Milan, et al.
Published: (2022)
by: Češka, Milan, et al.
Published: (2022)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
Algorithmic Strategies for Sustainable Reuse of Neural Network Accelerators with Permanent Faults
by: Alama, Youssef A. Ait, et al.
Published: (2024)
by: Alama, Youssef A. Ait, et al.
Published: (2024)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
by: Fang, Jiaxun, et al.
Published: (2025)
by: Fang, Jiaxun, et al.
Published: (2025)
PhotoGAN: Generative Adversarial Neural Network Acceleration with Silicon Photonics
by: Suresh, Tharini, et al.
Published: (2025)
by: Suresh, Tharini, et al.
Published: (2025)
A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms
by: Silvano, Cristina, et al.
Published: (2023)
by: Silvano, Cristina, et al.
Published: (2023)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
by: Xiang, Maoyang, et al.
Published: (2025)
by: Xiang, Maoyang, et al.
Published: (2025)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
by: Wang, Irene, et al.
Published: (2025)
by: Wang, Irene, et al.
Published: (2025)
Similar Items
-
AxMED: Formal Analysis and Automated Design of Approximate Median Filters using BDDs
by: Mrazek, Vojtech, et al.
Published: (2025) -
TRAPTI: Time-Resolved Analysis for SRAM Banking and Power Gating Optimization in Embedded Transformer Inference
by: Klhufek, Jan, et al.
Published: (2026) -
Late Breaking Result: FPGA-Based Emulation and Fault Injection for CNN Inference Accelerators
by: Masar, Filip, et al.
Published: (2025) -
ObfAx: Obfuscation and IP Piracy Detection in Approximate Circuits
by: Sekanina, Lukas, et al.
Published: (2026) -
Evolutionary Approximation of Ternary Neurons for On-sensor Printed Neural Networks
by: Mrazek, Vojtech, et al.
Published: (2024)