da4ml: Distributed Arithmetic for Real-time Neural Networks on FPGAs
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Chang, Que, Zhiqiang, Loncar, Vladimir, Luk, Wayne, Spiropulu, Maria |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Approximate Logic Synthesis Using BLASYS
by: Ma, Jingxiao, et al.
Published: (2025)
by: Ma, Jingxiao, et al.
Published: (2025)
JEDI-linear: Fast and Efficient Graph Neural Networks for Jet Tagging on FPGAs
by: Que, Zhiqiang, et al.
Published: (2025)
by: Que, Zhiqiang, et al.
Published: (2025)
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
by: Sun, Chang, et al.
Published: (2026)
by: Sun, Chang, et al.
Published: (2026)
LLM-based Behaviour Driven Development for Hardware Design
by: Drechsler, Rolf, et al.
Published: (2025)
by: Drechsler, Rolf, et al.
Published: (2025)
NOVA: NoC-based Vector Unit for Mapping Attention Layers on a CNN Accelerator
by: Upadhyay, Mohit, et al.
Published: (2024)
by: Upadhyay, Mohit, et al.
Published: (2024)
LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy Physics
by: Que, Zhiqiang, et al.
Published: (2022)
by: Que, Zhiqiang, et al.
Published: (2022)
Multiplier-free In-Memory Vector-Matrix Multiplication Using Distributed Arithmetic
by: Zeller, Felix, et al.
Published: (2025)
by: Zeller, Felix, et al.
Published: (2025)
Veryl: A New Hardware Description Language as an Altarnative to SystemVerilog
by: Hatta, Naoya, et al.
Published: (2024)
by: Hatta, Naoya, et al.
Published: (2024)
Memory Access Vectors: Improving Sampling Fidelity for CPU Performance Simulations
by: Caculo, Sriyash, et al.
Published: (2025)
by: Caculo, Sriyash, et al.
Published: (2025)
SparsePixels: Efficient Convolution for Sparse Data on FPGAs
by: Tsoi, Ho Fung, et al.
Published: (2025)
by: Tsoi, Ho Fung, et al.
Published: (2025)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
Photonic AI: A Hybrid Diffractive Holographic Neural System for Passive Optical Real-Time Image Classification
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
Comprehensive Formal Verification of Observational Correctness for the CHERIoT-Ibex Processor
by: Ploix, Louis-Emile, et al.
Published: (2025)
by: Ploix, Louis-Emile, et al.
Published: (2025)
Finite-Time Lyapunov Exponent Calculation on FPGA using High-Level Synthesis Tools
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
Trust Nothing: RTOS Security without Run-Time Software TCB (Extended Version)
by: Ackermann, Eric, et al.
Published: (2026)
by: Ackermann, Eric, et al.
Published: (2026)
Re-thinking Memory-Bound Limitations in CGRAs
by: Liu, Xiangfeng, et al.
Published: (2025)
by: Liu, Xiangfeng, et al.
Published: (2025)
AxOCS: Scaling FPGA-based Approximate Operators using Configuration Supersampling
by: Sahoo, Siva Satyendra, et al.
Published: (2023)
by: Sahoo, Siva Satyendra, et al.
Published: (2023)
Low-Latency FPGA Control System for Real-Time Neural Network Processing in CCD-Based Trapped-Ion Qubit Measurement
by: Lou, Binglei, et al.
Published: (2025)
by: Lou, Binglei, et al.
Published: (2025)
An SMT Formalization of Mixed-Precision Matrix Multiplication: Modeling Three Generations of Tensor Cores
by: Valpey, Benjamin, et al.
Published: (2025)
by: Valpey, Benjamin, et al.
Published: (2025)
CLIPGen: A Chiplet Link IP Modeling and Generation Framework for 2.5D Architecture Exploration
by: Zhu, Zhengping, et al.
Published: (2026)
by: Zhu, Zhengping, et al.
Published: (2026)
FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
by: Ramhorst, Benjamin, et al.
Published: (2023)
by: Ramhorst, Benjamin, et al.
Published: (2023)
hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware
by: Schulte, Jan-Frederik, et al.
Published: (2025)
by: Schulte, Jan-Frederik, et al.
Published: (2025)
AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
by: Danopoulos, Dimitrios, et al.
Published: (2025)
by: Danopoulos, Dimitrios, et al.
Published: (2025)
MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration
by: Que, Zhiqiang, et al.
Published: (2025)
by: Que, Zhiqiang, et al.
Published: (2025)
Regular mixed-radix DFT matrix factorization for in-place FFT accelerators
by: Salishev, Sergey
Published: (2025)
by: Salishev, Sergey
Published: (2025)
An Integrated UVM-TLM Co-Simulation Framework for RISC-V Functional Verification and Performance Evaluation
by: Qiu, Ruizhi, et al.
Published: (2025)
by: Qiu, Ruizhi, et al.
Published: (2025)
Ultra Fast Transformers on FPGAs for Particle Physics Experiments
by: Jiang, Zhixing, et al.
Published: (2024)
by: Jiang, Zhixing, et al.
Published: (2024)
Transaction Level Hierarchy Guided and Functional Coverage Driven Deductive Formal Verification
by: Strauch, Tobias
Published: (2025)
by: Strauch, Tobias
Published: (2025)
HGQ: High Granularity Quantization for Real-time Neural Networks on FPGAs
by: Sun, Chang, et al.
Published: (2024)
by: Sun, Chang, et al.
Published: (2024)
ALL-MASK: A Reconfigurable Logic Locking Method for Multicore Architecture with Sequential-Instruction-Oriented Key
by: Wang, Jianfeng, et al.
Published: (2022)
by: Wang, Jianfeng, et al.
Published: (2022)
SR-NCL: an Area-/Energy-Efficient Resilient NCL Architecture Based on Selective Redundancy
by: Ziad, Hasnain A., et al.
Published: (2025)
by: Ziad, Hasnain A., et al.
Published: (2025)
Integrating SystemC-AMS Power Modeling with a RISC-V ISS for Virtual Prototyping of Battery-operated Embedded Devices
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
An Improved Template for Approximate Computing
by: Rezaalipour, Morteza, et al.
Published: (2025)
by: Rezaalipour, Morteza, et al.
Published: (2025)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
by: Aspros, Maxime Henri, et al.
Published: (2025)
by: Aspros, Maxime Henri, et al.
Published: (2025)
Towards Automated Verification of Logarithmic Arithmetic
by: Arnold, Mark G., et al.
Published: (2024)
by: Arnold, Mark G., et al.
Published: (2024)
Real-Time Stream Compaction for Sparse Machine Learning on FPGAs
by: Neu, Marc, et al.
Published: (2026)
by: Neu, Marc, et al.
Published: (2026)
Enthuse: Efficient Adaptable High-throughput Streaming Aggregation Engines
by: Papaphilippou, Philippos, et al.
Published: (2024)
by: Papaphilippou, Philippos, et al.
Published: (2024)
ECOLogic: Enabling Circular, Obfuscated, and Adaptive Logic via eFPGA-Augmented SoCs
by: Tashdid, Ishraq, et al.
Published: (2025)
by: Tashdid, Ishraq, et al.
Published: (2025)
Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
by: Li, Richie
Published: (2025)
by: Li, Richie
Published: (2025)
Similar Items
-
Approximate Logic Synthesis Using BLASYS
by: Ma, Jingxiao, et al.
Published: (2025) -
JEDI-linear: Fast and Efficient Graph Neural Networks for Jet Tagging on FPGAs
by: Que, Zhiqiang, et al.
Published: (2025) -
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
by: Sun, Chang, et al.
Published: (2026) -
LLM-based Behaviour Driven Development for Hardware Design
by: Drechsler, Rolf, et al.
Published: (2025) -
NOVA: NoC-based Vector Unit for Mapping Attention Layers on a CNN Accelerator
by: Upadhyay, Mohit, et al.
Published: (2024)