Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
Fuente:
arXiv
Saved in:
| Main Authors: | Lübeck, Konstantin, Jung, Alexander Louis-Ferdinand, Wedlich, Felix, Müller, Mika Markus, Peccia, Federico Nicolás, Thömmes, Felix, Steinmetz, Jannik, Biermaier, Valentin, Frischknecht, Adrian, Bernardo, Paul Palomero, Bringmann, Oliver |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025)
by: Müller, Mika Markus, et al.
Published: (2025)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
Using the Abstract Computer Architecture Description Language to Model AI Hardware Accelerators
by: Müller, Mika Markus, et al.
Published: (2024)
by: Müller, Mika Markus, et al.
Published: (2024)
A Configurable and Efficient Memory Hierarchy for Neural Network Hardware Accelerator
by: Bause, Oliver, et al.
Published: (2024)
by: Bause, Oliver, et al.
Published: (2024)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
by: Peccia, Federico Nicolas, et al.
Published: (2024)
by: Peccia, Federico Nicolas, et al.
Published: (2024)
Smart Video Capsule Endoscopy: Raw Image-Based Localization for Enhanced GI Tract Investigation
by: Bause, Oliver, et al.
Published: (2025)
by: Bause, Oliver, et al.
Published: (2025)
HAPM -- Hardware Aware Pruning Method for CNN hardware accelerators in resource constrained devices
by: Peccia, Federico Nicolas, et al.
Published: (2024)
by: Peccia, Federico Nicolas, et al.
Published: (2024)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
by: Wang, Aotao, et al.
Published: (2025)
by: Wang, Aotao, et al.
Published: (2025)
Stream-HLS: Towards Automatic Dataflow Acceleration
by: Basalama, Suhail, et al.
Published: (2025)
by: Basalama, Suhail, et al.
Published: (2025)
Efficient yet Accurate End-to-End SC Accelerator Design
by: Li, Meng, et al.
Published: (2024)
by: Li, Meng, et al.
Published: (2024)
ATLAAS: Automatic Tensor-Level Abstraction of Accelerator Semantics
by: Gao, Ruijie, et al.
Published: (2026)
by: Gao, Ruijie, et al.
Published: (2026)
FireBridge: Cycle-Accurate Hardware + Firmware Co-Verification for Modern Accelerators
by: Abarajithan, G, et al.
Published: (2026)
by: Abarajithan, G, et al.
Published: (2026)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
by: Lee, Kyungmi, et al.
Published: (2026)
by: Lee, Kyungmi, et al.
Published: (2026)
CiMLoop: A Flexible, Accurate, and Fast Compute-In-Memory Modeling Tool
by: Andrulis, Tanner, et al.
Published: (2024)
by: Andrulis, Tanner, et al.
Published: (2024)
Towards Efficient and Accurate Detection of On-Chip Fail-Slow Failures for Many-Core Accelerators
by: Wu, Junchi, et al.
Published: (2025)
by: Wu, Junchi, et al.
Published: (2025)
ZynqParrot: A Scale-Down Approach to Cycle-Accurate, FPGA-Accelerated Co-Emulation
by: Ruelas-Petrisko, Daniel, et al.
Published: (2025)
by: Ruelas-Petrisko, Daniel, et al.
Published: (2025)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
by: Yang, Kuilian, et al.
Published: (2026)
by: Yang, Kuilian, et al.
Published: (2026)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
by: Wang, Xuan, et al.
Published: (2024)
by: Wang, Xuan, et al.
Published: (2024)
EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques
by: Canpolat, Oğuzhan, et al.
Published: (2025)
by: Canpolat, Oğuzhan, et al.
Published: (2025)
PAI: Fast, Accurate, and Full Benchmark Performance Projection with AI
by: Johnson, Avery, et al.
Published: (2026)
by: Johnson, Avery, et al.
Published: (2026)
Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal
by: Jiang, Wenqi, et al.
Published: (2024)
by: Jiang, Wenqi, et al.
Published: (2024)
From PyTorch to Calyx: An Open-Source Compiler Toolchain for ML Accelerators
by: Xie, Jiahan, et al.
Published: (2025)
by: Xie, Jiahan, et al.
Published: (2025)
GateKeeper-GPU: Fast and Accurate Pre-Alignment Filtering in Short Read Mapping
by: Bingöl, Zülal, et al.
Published: (2021)
by: Bingöl, Zülal, et al.
Published: (2021)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
by: Ding, Hong, et al.
Published: (2025)
by: Ding, Hong, et al.
Published: (2025)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
by: Umuroglu, Yaman, et al.
Published: (2025)
by: Umuroglu, Yaman, et al.
Published: (2025)
FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
by: Rahoof, Abdul, et al.
Published: (2025)
by: Rahoof, Abdul, et al.
Published: (2025)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
by: Zhang, Junming, et al.
Published: (2026)
by: Zhang, Junming, et al.
Published: (2026)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
by: Qararyah, Fareed, et al.
Published: (2025)
by: Qararyah, Fareed, et al.
Published: (2025)
ACT: Automatically Generating Compiler Backends from Tensor Accelerator ISA Descriptions
by: Jain, Devansh, et al.
Published: (2025)
by: Jain, Devansh, et al.
Published: (2025)
Fast and Fusiest: An Optimal Fusion-Aware Mapper for Accelerator Design
by: Andrulis, Tanner, et al.
Published: (2026)
by: Andrulis, Tanner, et al.
Published: (2026)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
by: Wadle, Shayne, et al.
Published: (2025)
by: Wadle, Shayne, et al.
Published: (2025)
Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation Methodology
by: Kanellopoulos, Konstantinos, et al.
Published: (2024)
by: Kanellopoulos, Konstantinos, et al.
Published: (2024)
Efficient and Accurate Graph Classification with Hyperdimensional Computing on FPGA
by: Arockiaraj, Jebacyril, et al.
Published: (2025)
by: Arockiaraj, Jebacyril, et al.
Published: (2025)
Work-in-Progress: Real-Time Neural Network Inference on a Custom RISC-V Multicore Vector Processor
by: Kirschner, Maximilian, et al.
Published: (2024)
by: Kirschner, Maximilian, et al.
Published: (2024)
BOLT: Bandwidth-Optimized Lightning-Fast Oblivious Map powered by Secure HBM Accelerators
by: Guo, Yitong, et al.
Published: (2025)
by: Guo, Yitong, et al.
Published: (2025)
Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
by: Agosta, Giovanni, et al.
Published: (2025)
by: Agosta, Giovanni, et al.
Published: (2025)
The Turbo-Charged Mapper: Fast and Optimal Mapping for Energy-efficient and Low-latency Accelerator Design
by: Gilbert, Michael, et al.
Published: (2026)
by: Gilbert, Michael, et al.
Published: (2026)
CiFHER: A Chiplet-Based FHE Accelerator with a Resizable Structure
by: Kim, Sangpyo, et al.
Published: (2023)
by: Kim, Sangpyo, et al.
Published: (2023)
CLAASIC: a Cortex-Inspired Hardware Accelerator
by: Puente, Valentin, et al.
Published: (2016)
by: Puente, Valentin, et al.
Published: (2016)
Similar Items
-
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025) -
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024) -
Using the Abstract Computer Architecture Description Language to Model AI Hardware Accelerators
by: Müller, Mika Markus, et al.
Published: (2024) -
A Configurable and Efficient Memory Hierarchy for Neural Network Hardware Accelerator
by: Bause, Oliver, et al.
Published: (2024) -
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
by: Peccia, Federico Nicolas, et al.
Published: (2024)