Performance Analysis of DNN Inference/Training with Convolution and non-Convolution Operations
Fuente:
arXiv
Saved in:
| Main Authors: | Esmaeilzadeh, Hadi, Ghodrati, Soroush, Kahng, Andrew B., Kinzer, Sean, Manasi, Susmita Dey, Sapatnekar, Sachin S., Wang, Zhiang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Invited: Toward Accurate, Large-scale Electromigration Analysis and Optimization in Integrated Systems
by: Sapatnekar, Sachin S.
Published: (2026)
by: Sapatnekar, Sachin S.
Published: (2026)
DG-RePlAce: A Dataflow-Driven GPU-Accelerated Analytical Global Placement Framework for Machine Learning Accelerators
by: Kahng, Andrew B., et al.
Published: (2024)
by: Kahng, Andrew B., et al.
Published: (2024)
In-Storage Domain-Specific Acceleration for Serverless Computing
by: Mahapatra, Rohan, et al.
Published: (2023)
by: Mahapatra, Rohan, et al.
Published: (2023)
Accelerating Electrostatics-based Global Placement with Enhanced FFT Computation
by: Zhang, Hangyu, et al.
Published: (2025)
by: Zhang, Hangyu, et al.
Published: (2025)
ChipletPart: Cost-Aware Partitioning for 2.5D Systems
by: Graening, Alexander, et al.
Published: (2025)
by: Graening, Alexander, et al.
Published: (2025)
Invited: Toward Sustainable and Transparent Benchmarking for Academic Physical Design Research
by: Jiang, Liwen, et al.
Published: (2026)
by: Jiang, Liwen, et al.
Published: (2026)
ML-based AIG Timing Prediction to Enhance Logic Optimization
by: Jiang, Wenjing, et al.
Published: (2024)
by: Jiang, Wenjing, et al.
Published: (2024)
A Linear-Time Algorithm for Steady-State Analysis of Electromigration in General Interconnects
by: Shohel, Mohammad Abdullah Al, et al.
Published: (2021)
by: Shohel, Mohammad Abdullah Al, et al.
Published: (2021)
IR-Aware ECO Timing Optimization Using Reinforcement Learning
by: Jiang, Wenjing, et al.
Published: (2024)
by: Jiang, Wenjing, et al.
Published: (2024)
COmPOSER: Circuit Optimization of mm-wave/RF circuits with Performance-Oriented Synthesis for Efficient Realizations
by: Ghosh, Subhadip, et al.
Published: (2026)
by: Ghosh, Subhadip, et al.
Published: (2026)
ECO-CHIP: Estimation of Carbon Footprint of Chiplet-based Architectures for Sustainable VLSI
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
Tiny Chiplets Enabled by Packaging Scaling: Opportunities in ESD Protection and Signal Integrity
by: Haque, Emad, et al.
Published: (2025)
by: Haque, Emad, et al.
Published: (2025)
Escaping Flatland: A Placement Flow for Enabling 3D FPGAs
by: Hao, Cong, et al.
Published: (2026)
by: Hao, Cong, et al.
Published: (2026)
High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
by: Lin, Kuan-Ting, et al.
Published: (2025)
by: Lin, Kuan-Ting, et al.
Published: (2025)
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
An Extended Study of Gear-Ratio-Aware Standard Cell Layout Generation for DTCO Exploration
by: Cheng, Chung-Kuan, et al.
Published: (2026)
by: Cheng, Chung-Kuan, et al.
Published: (2026)
Convolutions Predictable Offloading to an Accelerator: Formalization and Optimization
by: Husson, Benjamin, et al.
Published: (2026)
by: Husson, Benjamin, et al.
Published: (2026)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
by: Sonnino, Lorenzo, et al.
Published: (2023)
by: Sonnino, Lorenzo, et al.
Published: (2023)
NeuroBlend: Towards Low-Power yet Accurate Neural Network-Based Inference Engine Blending Binary and Fixed-Point Convolutions
by: Fayyazi, Arash, et al.
Published: (2023)
by: Fayyazi, Arash, et al.
Published: (2023)
CADC: Crossbar-Aware Dendritic Convolution for Efficient In-memory Computing
by: Dong, Shuai, et al.
Published: (2025)
by: Dong, Shuai, et al.
Published: (2025)
A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
by: Ali, Sami Ben, et al.
Published: (2024)
by: Ali, Sami Ben, et al.
Published: (2024)
ArtNet: Hierarchical Clustering-Based Artificial Netlist Generator for ML and DTCO Application
by: Kang, Andrew B. Kahng. Seokhyeong, et al.
Published: (2025)
by: Kang, Andrew B. Kahng. Seokhyeong, et al.
Published: (2025)
Energy-Efficient FPGA Framework for Non-Quantized Convolutional Neural Networks
by: Athanasiadis, Angelos, et al.
Published: (2025)
by: Athanasiadis, Angelos, et al.
Published: (2025)
Leveraging Highly Approximated Multipliers in DNN Inference
by: Zervakis, Georgios, et al.
Published: (2024)
by: Zervakis, Georgios, et al.
Published: (2024)
SPEED: A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
Accelerating OTA Circuit Design: Transistor Sizing Based on a Transformer Model and Precomputed Lookup Tables
by: Ghosh, Subhadip, et al.
Published: (2025)
by: Ghosh, Subhadip, et al.
Published: (2025)
PULSE: Parametric Hardware Units for Low-power Sparsity-Aware Convolution Engine
by: Aliyev, Ilkin, et al.
Published: (2024)
by: Aliyev, Ilkin, et al.
Published: (2024)
PowerFlow-DNN: Compiler-Directed Fine-Grained Power Orchestration for End-to-End Edge AI Inference
by: Chen, Paul, et al.
Published: (2026)
by: Chen, Paul, et al.
Published: (2026)
Fair and Square: Replacing One Real Multiplication with a Single Square and One Complex Multiplication with Three Squares When Performing Matrix Multiplication and Convolutions
by: Liguori, Vincenzo
Published: (2026)
by: Liguori, Vincenzo
Published: (2026)
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
by: Sestito, Cristian, et al.
Published: (2025)
by: Sestito, Cristian, et al.
Published: (2025)
Demystifying the 7-D Convolution Loop Nest for Data and Instruction Streaming in Reconfigurable AI Accelerators
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
by: Cherezova, Natalia, et al.
Published: (2025)
by: Cherezova, Natalia, et al.
Published: (2025)
TetrisG-SDK: Efficient Convolutional Layer Mapping with Adaptive Windows and Grouped Convolutions for Fast In-Memory Computing
by: Dong, Ke, et al.
Published: (2026)
by: Dong, Ke, et al.
Published: (2026)
Hardware-Aware DNN Compression for Homogeneous Edge Devices
by: Zhang, Kunlong, et al.
Published: (2025)
by: Zhang, Kunlong, et al.
Published: (2025)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
by: Chen, Xingzhen, et al.
Published: (2026)
by: Chen, Xingzhen, et al.
Published: (2026)
A Novel FPGA-based CNN Hardware Accelerator: Optimization for Convolutional Layers using Karatsuba Ofman Multiplier
by: Sarkar, Amit
Published: (2024)
by: Sarkar, Amit
Published: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
by: Sun, Xiaotian, et al.
Published: (2024)
by: Sun, Xiaotian, et al.
Published: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024)
by: Qararyah, Fareed, et al.
Published: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
Similar Items
-
Invited: Toward Accurate, Large-scale Electromigration Analysis and Optimization in Integrated Systems
by: Sapatnekar, Sachin S.
Published: (2026) -
DG-RePlAce: A Dataflow-Driven GPU-Accelerated Analytical Global Placement Framework for Machine Learning Accelerators
by: Kahng, Andrew B., et al.
Published: (2024) -
In-Storage Domain-Specific Acceleration for Serverless Computing
by: Mahapatra, Rohan, et al.
Published: (2023) -
Accelerating Electrostatics-based Global Placement with Enhanced FFT Computation
by: Zhang, Hangyu, et al.
Published: (2025) -
ChipletPart: Cost-Aware Partitioning for 2.5D Systems
by: Graening, Alexander, et al.
Published: (2025)