Combining Power and Arithmetic Optimization via Datapath Rewriting
Fuente:
arXiv
Salvato in:
| Autori principali: | Coward, Samuel, Drane, Theo, Morini, Emiliano, Constantinides, George |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ROVER: RTL Optimization via Verified E-Graph Rewriting
di: Coward, Samuel, et al.
Pubblicazione: (2024)
di: Coward, Samuel, et al.
Pubblicazione: (2024)
On the Systematic Creation of Faithfully Rounded Commutative Truncated Booth Multipliers
di: Drane, Theo, et al.
Pubblicazione: (2024)
di: Drane, Theo, et al.
Pubblicazione: (2024)
ReducedLUT: Table Decomposition with "Don't Care" Conditions
di: Cassidy, Oliver, et al.
Pubblicazione: (2024)
di: Cassidy, Oliver, et al.
Pubblicazione: (2024)
Refining Datapath for Microscaling ViTs
di: Xiao, Can, et al.
Pubblicazione: (2025)
di: Xiao, Can, et al.
Pubblicazione: (2025)
Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap
di: Langhammer, Martin, et al.
Pubblicazione: (2024)
di: Langhammer, Martin, et al.
Pubblicazione: (2024)
A Statically and Dynamically Scalable Soft GPGPU
di: Langhammer, Martin, et al.
Pubblicazione: (2024)
di: Langhammer, Martin, et al.
Pubblicazione: (2024)
Banked Memories for Soft SIMT Processors
di: Langhammer, Martin, et al.
Pubblicazione: (2025)
di: Langhammer, Martin, et al.
Pubblicazione: (2025)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
di: Chen, Zhihan, et al.
Pubblicazione: (2024)
di: Chen, Zhihan, et al.
Pubblicazione: (2024)
Designing Approximate Arithmetic Circuits with Combined Error Constraints
di: Češka, Milan, et al.
Pubblicazione: (2022)
di: Češka, Milan, et al.
Pubblicazione: (2022)
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
di: Hayes, Oran, et al.
Pubblicazione: (2026)
di: Hayes, Oran, et al.
Pubblicazione: (2026)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
di: Andronic, Marta, et al.
Pubblicazione: (2024)
di: Andronic, Marta, et al.
Pubblicazione: (2024)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
di: Andronic, Marta, et al.
Pubblicazione: (2023)
di: Andronic, Marta, et al.
Pubblicazione: (2023)
FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
di: Ramhorst, Benjamin, et al.
Pubblicazione: (2023)
di: Ramhorst, Benjamin, et al.
Pubblicazione: (2023)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
di: Andronic, Marta, et al.
Pubblicazione: (2025)
di: Andronic, Marta, et al.
Pubblicazione: (2025)
High-Performance Pipelined NTT Accelerators with Homogeneous Digit-Serial Modulo Arithmetic
di: Alexakis, George, et al.
Pubblicazione: (2025)
di: Alexakis, George, et al.
Pubblicazione: (2025)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
di: Cheng, Jianyi, et al.
Pubblicazione: (2023)
di: Cheng, Jianyi, et al.
Pubblicazione: (2023)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
di: Biggs, Benjamin, et al.
Pubblicazione: (2023)
di: Biggs, Benjamin, et al.
Pubblicazione: (2023)
Exploring FPGA designs for MX and beyond
di: Samson, Ebby, et al.
Pubblicazione: (2024)
di: Samson, Ebby, et al.
Pubblicazione: (2024)
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
di: Cuyckens, Stef, et al.
Pubblicazione: (2025)
di: Cuyckens, Stef, et al.
Pubblicazione: (2025)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
Modulo-$(2^{2n}+1)$ Arithmetic via Two Parallel n-bit Residue Channels
di: Jaberipur, Ghassem, et al.
Pubblicazione: (2024)
di: Jaberipur, Ghassem, et al.
Pubblicazione: (2024)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
di: Wu, Jiajun, et al.
Pubblicazione: (2024)
di: Wu, Jiajun, et al.
Pubblicazione: (2024)
Increasing the Energy-Efficiency of Wearables Using Low-Precision Posit Arithmetic with PHEE
di: Mallasén, David, et al.
Pubblicazione: (2025)
di: Mallasén, David, et al.
Pubblicazione: (2025)
LLM-Powered Code Analysis and Optimization for Gaussian Splatting Kernels
di: Hu, Yi, et al.
Pubblicazione: (2025)
di: Hu, Yi, et al.
Pubblicazione: (2025)
Big-PERCIVAL: Exploring the Native Use of 64-Bit Posit Arithmetic in Scientific Computing
di: Mallasén, David, et al.
Pubblicazione: (2023)
di: Mallasén, David, et al.
Pubblicazione: (2023)
DSLR-CNN: Efficient CNN Acceleration using Digit-Serial Left-to-Right Arithmetic
di: Nisar, Malik Zohaib, et al.
Pubblicazione: (2025)
di: Nisar, Malik Zohaib, et al.
Pubblicazione: (2025)
Combining Fault Tolerance Techniques and COTS SoC Accelerators for Payload Processing in Space
di: Leon, Vasileios, et al.
Pubblicazione: (2025)
di: Leon, Vasileios, et al.
Pubblicazione: (2025)
AC-Refiner: Efficient Arithmetic Circuit Optimization Using Conditional Diffusion Models
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
A Low-Power Sparse Deep Learning Accelerator with Optimized Data Reuse
di: Hsu, Kai-Chieh, et al.
Pubblicazione: (2025)
di: Hsu, Kai-Chieh, et al.
Pubblicazione: (2025)
From Circuits to SoC Processors: Arithmetic Approximation Techniques & Embedded Computing Methodologies for DSP Acceleration
di: Leon, Vasileios
Pubblicazione: (2023)
di: Leon, Vasileios
Pubblicazione: (2023)
Optimizing Scalable Multi-Cluster Architectures for Next-Generation Wireless Sensing and Communication
di: Riedel, Samuel, et al.
Pubblicazione: (2025)
di: Riedel, Samuel, et al.
Pubblicazione: (2025)
TRAPTI: Time-Resolved Analysis for SRAM Banking and Power Gating Optimization in Embedded Transformer Inference
di: Klhufek, Jan, et al.
Pubblicazione: (2026)
di: Klhufek, Jan, et al.
Pubblicazione: (2026)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
di: Chen, Yuzong, et al.
Pubblicazione: (2024)
di: Chen, Yuzong, et al.
Pubblicazione: (2024)
AutoPower: Automated Few-Shot Architecture-Level Power Modeling by Power Group Decoupling
di: Zhang, Qijun, et al.
Pubblicazione: (2025)
di: Zhang, Qijun, et al.
Pubblicazione: (2025)
FirePower: Towards a Foundation with Generalizable Knowledge for Architecture-Level Power Modeling
di: Zhang, Qijun, et al.
Pubblicazione: (2024)
di: Zhang, Qijun, et al.
Pubblicazione: (2024)
Soft Error Probability Estimation of Nano-scale Combinational Circuits
di: Jockar, Ali, et al.
Pubblicazione: (2025)
di: Jockar, Ali, et al.
Pubblicazione: (2025)
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
di: Wu, Qizhe, et al.
Pubblicazione: (2024)
di: Wu, Qizhe, et al.
Pubblicazione: (2024)
ReadyPower: A Reliable, Interpretable, and Handy Architectural Power Model Based on Analytical Framework
di: Zhang, Qijun, et al.
Pubblicazione: (2025)
di: Zhang, Qijun, et al.
Pubblicazione: (2025)
E-Syn: E-Graph Rewriting with Technology-Aware Cost Functions for Logic Synthesis
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
Attacking AI Accelerators by Leveraging Arithmetic Properties of Addition
di: Heidary, Masoud, et al.
Pubblicazione: (2026)
di: Heidary, Masoud, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ROVER: RTL Optimization via Verified E-Graph Rewriting
di: Coward, Samuel, et al.
Pubblicazione: (2024) -
On the Systematic Creation of Faithfully Rounded Commutative Truncated Booth Multipliers
di: Drane, Theo, et al.
Pubblicazione: (2024) -
ReducedLUT: Table Decomposition with "Don't Care" Conditions
di: Cassidy, Oliver, et al.
Pubblicazione: (2024) -
Refining Datapath for Microscaling ViTs
di: Xiao, Can, et al.
Pubblicazione: (2025) -
Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap
di: Langhammer, Martin, et al.
Pubblicazione: (2024)