PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
Fuente:
arXiv
Saved in:
| Main Authors: | Saha, Rappy, Haris, Jude, Agostini, Nicolas Bohm, Kaeli, David, Cano, José |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating PoT Quantization on Edge Devices
by: Saha, Rappy, et al.
Published: (2024)
by: Saha, Rappy, et al.
Published: (2024)
Designing Efficient LLM Accelerators for Edge Devices
by: Haris, Jude, et al.
Published: (2024)
by: Haris, Jude, et al.
Published: (2024)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
by: Sharma, Vinamra, et al.
Published: (2026)
by: Sharma, Vinamra, et al.
Published: (2026)
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
by: Yang, Hanchen, et al.
Published: (2025)
by: Yang, Hanchen, et al.
Published: (2025)
Accelerating Transposed Convolutions on FPGA-based Edge Devices
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators
by: Agostini, Nicolas Bohm, et al.
Published: (2023)
by: Agostini, Nicolas Bohm, et al.
Published: (2023)
A2Q+: Improving Accumulator-Aware Weight Quantization
by: Colbert, Ian, et al.
Published: (2024)
by: Colbert, Ian, et al.
Published: (2024)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
by: Zhou, Cyrus, et al.
Published: (2023)
by: Zhou, Cyrus, et al.
Published: (2023)
PPU: Design and Implementation of a Pipelined Full Posit Processing Unit
by: Rossi, Federico, et al.
Published: (2023)
by: Rossi, Federico, et al.
Published: (2023)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
by: Chhugani, Jatin, et al.
Published: (2026)
by: Chhugani, Jatin, et al.
Published: (2026)
A Review on Proprietary Accelerators for Large Language Models
by: Park, Sihyeong, et al.
Published: (2025)
by: Park, Sihyeong, et al.
Published: (2025)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
by: Qararyah, Fareed, et al.
Published: (2025)
by: Qararyah, Fareed, et al.
Published: (2025)
AI Load Dynamics--A Power Electronics Perspective
by: Li, Yuzhuo, et al.
Published: (2025)
by: Li, Yuzhuo, et al.
Published: (2025)
Accelerating Transistor-Level Simulation of Integrated Circuits via Equivalence of RC Long-Chain Structures
by: Tang, Ruibai, et al.
Published: (2025)
by: Tang, Ruibai, et al.
Published: (2025)
A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
by: Kuper, Reese, et al.
Published: (2023)
by: Kuper, Reese, et al.
Published: (2023)
Makinote: An FPGA-Based HW/SW Platform for Pre-Silicon Emulation of RISC-V Designs
by: Perdomo, Elias, et al.
Published: (2024)
by: Perdomo, Elias, et al.
Published: (2024)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
by: Mazzola, Sergio, et al.
Published: (2025)
by: Mazzola, Sergio, et al.
Published: (2025)
Graph neural networks with configuration cross-attention for tensor compilers
by: Khizbullin, Dmitrii, et al.
Published: (2024)
by: Khizbullin, Dmitrii, et al.
Published: (2024)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
by: Lübeck, Konstantin, et al.
Published: (2024)
by: Lübeck, Konstantin, et al.
Published: (2024)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
GeneTEK: Low-power, high-performance and scalable FPGA architecture for exact unit-cost edit distance
by: Espinosa, Elena, et al.
Published: (2025)
by: Espinosa, Elena, et al.
Published: (2025)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
by: Ibrahim, Muhammad Sohail, et al.
Published: (2024)
by: Ibrahim, Muhammad Sohail, et al.
Published: (2024)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
by: Zhang, Tim, et al.
Published: (2022)
by: Zhang, Tim, et al.
Published: (2022)
Search Your Block Floating Point Scales!
by: Gupta, Tanmaey, et al.
Published: (2026)
by: Gupta, Tanmaey, et al.
Published: (2026)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
by: Atmer, Hannah, et al.
Published: (2025)
by: Atmer, Hannah, et al.
Published: (2025)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
by: Yang, Xiaoxuan, et al.
Published: (2025)
by: Yang, Xiaoxuan, et al.
Published: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025)
by: Müller, Mika Markus, et al.
Published: (2025)
The Bicameral Cache: a split cache for vector architectures
by: Rebolledo, Susana, et al.
Published: (2024)
by: Rebolledo, Susana, et al.
Published: (2024)
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
by: Tajdari, Sabiha, et al.
Published: (2025)
by: Tajdari, Sabiha, et al.
Published: (2025)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
by: Chen, Yao, et al.
Published: (2022)
by: Chen, Yao, et al.
Published: (2022)
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
by: Shivdikar, Kaustubh, et al.
Published: (2024)
by: Shivdikar, Kaustubh, et al.
Published: (2024)
ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
by: Liu, Hongxiang, et al.
Published: (2025)
by: Liu, Hongxiang, et al.
Published: (2025)
Heterogeneous Memory Benchmarking Toolkit
by: Ghaemi, Golsana, et al.
Published: (2025)
by: Ghaemi, Golsana, et al.
Published: (2025)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
by: Ali, Wajid, et al.
Published: (2025)
by: Ali, Wajid, et al.
Published: (2025)
Similar Items
-
Accelerating PoT Quantization on Edge Devices
by: Saha, Rappy, et al.
Published: (2024) -
Designing Efficient LLM Accelerators for Edge Devices
by: Haris, Jude, et al.
Published: (2024) -
LLM-Driven Design Space Exploration of FPGA-based Accelerators
by: Sharma, Vinamra, et al.
Published: (2026) -
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
by: Haris, Jude, et al.
Published: (2025) -
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
by: Yang, Hanchen, et al.
Published: (2025)