NeuroBlend: Towards Low-Power yet Accurate Neural Network-Based Inference Engine Blending Binary and Fixed-Point Convolutions
Fuente:
arXiv
Saved in:
| Main Authors: | Fayyazi, Arash, Nazemi, Mahdi, Fayyazi, Arya, Pedram, Massoud |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Co-Optimization Compiler: Leveraging Multi-Agent Reinforcement Learning for Enhanced DNN Accelerator Performance
by: Fayyazi, Arya, et al.
Published: (2024)
by: Fayyazi, Arya, et al.
Published: (2024)
CHOSEN: Compilation to Hardware Optimization Stack for Efficient Vision Transformer Inference
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
Sensitivity-Aware Mixed-Precision Quantization and Width Optimization of Deep Neural Networks Through Cluster-Based Tree-Structured Parzen Estimation
by: Azizi, Seyedarmin, et al.
Published: (2023)
by: Azizi, Seyedarmin, et al.
Published: (2023)
MENAGE: Mixed-Signal Event-Driven Neuromorphic Accelerator for Edge Applications
by: Abdollahi, Armin, et al.
Published: (2024)
by: Abdollahi, Armin, et al.
Published: (2024)
IC-D2S: A Hybrid Ising-Classical-Machines Data-Driven QUBO Solver Method
by: Abdollahi, Armin, et al.
Published: (2025)
by: Abdollahi, Armin, et al.
Published: (2025)
Reuse and Blend: Energy-Efficient Optical Neural Network Enabled by Weight Sharing
by: Xu, Bo, et al.
Published: (2024)
by: Xu, Bo, et al.
Published: (2024)
SAIM: Scalable Analog Ising Machine for Solving Quadratic Binary Optimization Problems
by: Razmkhah, Sasan, et al.
Published: (2024)
by: Razmkhah, Sasan, et al.
Published: (2024)
HDLFORGE: A Two-Stage Multi-Agent Framework for Efficient Verilog Code Generation with Adaptive Model Escalation
by: Abdollahi, Armin, et al.
Published: (2026)
by: Abdollahi, Armin, et al.
Published: (2026)
PULSE: Parametric Hardware Units for Low-power Sparsity-Aware Convolution Engine
by: Aliyev, Ilkin, et al.
Published: (2024)
by: Aliyev, Ilkin, et al.
Published: (2024)
Efficient yet Accurate End-to-End SC Accelerator Design
by: Li, Meng, et al.
Published: (2024)
by: Li, Meng, et al.
Published: (2024)
Dedicated Inference Engine and Binary-Weight Neural Networks for Lightweight Instance Segmentation
by: Chen, Tse-Wei, et al.
Published: (2025)
by: Chen, Tse-Wei, et al.
Published: (2025)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
by: Lin, Kuan-Ting, et al.
Published: (2025)
by: Lin, Kuan-Ting, et al.
Published: (2025)
TimingLLM: A Two-Stage Retrieval-Augmented Framework for Pre-Synthesis Timing Prediction from Verilog
by: Abdollahi, Armin, et al.
Published: (2026)
by: Abdollahi, Armin, et al.
Published: (2026)
Neuro-Photonix: Enabling Near-Sensor Neuro-Symbolic AI Computing on Silicon Photonics Substrate
by: Najafi, Deniz, et al.
Published: (2024)
by: Najafi, Deniz, et al.
Published: (2024)
Binary Neural Network Implementation for Handwritten Digit Recognition on FPGA
by: Ertörer, Emir Devlet, et al.
Published: (2025)
by: Ertörer, Emir Devlet, et al.
Published: (2025)
Energy-Efficient FPGA Framework for Non-Quantized Convolutional Neural Networks
by: Athanasiadis, Angelos, et al.
Published: (2025)
by: Athanasiadis, Angelos, et al.
Published: (2025)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)
by: Li, Haomin, et al.
Published: (2026)
Distributed Inference with Minimal Off-Chip Traffic for Transformers on Low-Power MCUs
by: Bochem, Severin, et al.
Published: (2024)
by: Bochem, Severin, et al.
Published: (2024)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
Gaussian Blending Unit: An Edge GPU Plug-in for Real-Time Gaussian-Based Rendering in AR/VR
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
MARCO: Hardware-Aware Neural Architecture Search for Edge Devices with Multi-Agent Reinforcement Learning and Conformal Prediction Filtering
by: Fayyazi, Arya, et al.
Published: (2025)
by: Fayyazi, Arya, et al.
Published: (2025)
Accelerating Elliptic Curve Point Additions on Versal AI Engine for Multi-scalar Multiplication
by: Ohno, Ayumi, et al.
Published: (2025)
by: Ohno, Ayumi, et al.
Published: (2025)
OPIMA: Optical Processing-In-Memory for Convolutional Neural Network Acceleration
by: Sunny, Febin, et al.
Published: (2024)
by: Sunny, Febin, et al.
Published: (2024)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
ASAP-FE: Energy-Efficient Feature Extraction Enabling Multi-Channel Keyword Spotting on Edge Processors
by: Choi, Jongin, et al.
Published: (2025)
by: Choi, Jongin, et al.
Published: (2025)
ONNX-to-Hardware Design Flow for Adaptive Neural-Network Inference on FPGAs
by: Manca, Federico, et al.
Published: (2024)
by: Manca, Federico, et al.
Published: (2024)
Ultra Low-Power SDM-based Circuit-Switching for Networks-on-Chip
by: Zaeemi, Meysam, et al.
Published: (2026)
by: Zaeemi, Meysam, et al.
Published: (2026)
NL-DPE: An Analog In-memory Non-Linear Dot Product Engine for Efficient CNN and LLM Inference
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
BinSparX: Sparsified Binary Neural Networks for Reduced Hardware Non-Idealities in Xbar Arrays
by: Malhotra, Akul, et al.
Published: (2024)
by: Malhotra, Akul, et al.
Published: (2024)
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
by: Natesh, Vikas, et al.
Published: (2025)
by: Natesh, Vikas, et al.
Published: (2025)
SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
by: Khabbazan, Bahareh, et al.
Published: (2025)
by: Khabbazan, Bahareh, et al.
Published: (2025)
Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review
by: Gareau, Jaël Champagne, et al.
Published: (2026)
by: Gareau, Jaël Champagne, et al.
Published: (2026)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
by: Lee, Kyungmi, et al.
Published: (2026)
by: Lee, Kyungmi, et al.
Published: (2026)
Voxel-CIM: An Efficient Compute-in-Memory Accelerator for Voxel-based Point Cloud Neural Networks
by: Lin, Xipeng, et al.
Published: (2024)
by: Lin, Xipeng, et al.
Published: (2024)
Dynamic Power Control in a Hardware Neural Network with Error-Configurable MAC Units
by: Ghaderi, Maedeh, et al.
Published: (2024)
by: Ghaderi, Maedeh, et al.
Published: (2024)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
Register Dispersion: Reducing the Footprint of the Vector Register File in Vector Engines of Low-Cost RISC-V CPUs
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Invited: Toward Accurate, Large-scale Electromigration Analysis and Optimization in Integrated Systems
by: Sapatnekar, Sachin S.
Published: (2026)
by: Sapatnekar, Sachin S.
Published: (2026)
Similar Items
-
Dynamic Co-Optimization Compiler: Leveraging Multi-Agent Reinforcement Learning for Enhanced DNN Accelerator Performance
by: Fayyazi, Arya, et al.
Published: (2024) -
CHOSEN: Compilation to Hardware Optimization Stack for Efficient Vision Transformer Inference
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024) -
Sensitivity-Aware Mixed-Precision Quantization and Width Optimization of Deep Neural Networks Through Cluster-Based Tree-Structured Parzen Estimation
by: Azizi, Seyedarmin, et al.
Published: (2023) -
MENAGE: Mixed-Signal Event-Driven Neuromorphic Accelerator for Edge Applications
by: Abdollahi, Armin, et al.
Published: (2024) -
IC-D2S: A Hybrid Ising-Classical-Machines Data-Driven QUBO Solver Method
by: Abdollahi, Armin, et al.
Published: (2025)