Fast and Fusiest: An Optimal Fusion-Aware Mapper for Accelerator Design
Fuente:
arXiv
Saved in:
| Main Authors: | Andrulis, Tanner, Gilbert, Michael, Sze, Vivienne, Emer, Joel S. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Turbo-Charged Mapper: Fast and Optimal Mapping for Energy-efficient and Low-latency Accelerator Design
by: Gilbert, Michael, et al.
Published: (2026)
by: Gilbert, Michael, et al.
Published: (2026)
CiMLoop: A Flexible, Accurate, and Fast Compute-In-Memory Modeling Tool
by: Andrulis, Tanner, et al.
Published: (2024)
by: Andrulis, Tanner, et al.
Published: (2024)
Modeling Analog-Digital-Converter Energy and Area for Compute-In-Memory Accelerator Design
by: Andrulis, Tanner, et al.
Published: (2024)
by: Andrulis, Tanner, et al.
Published: (2024)
Architecture-Level Modeling of Photonic Deep Neural Network Accelerators
by: Andrulis, Tanner, et al.
Published: (2024)
by: Andrulis, Tanner, et al.
Published: (2024)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
by: Gilbert, Michael, et al.
Published: (2024)
by: Gilbert, Michael, et al.
Published: (2024)
Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Capacity
by: Xue, Zi Yu, et al.
Published: (2023)
by: Xue, Zi Yu, et al.
Published: (2023)
FuseMax: Leveraging Extended Einsums to Optimize Attention Accelerator Design
by: Nayak, Nandeeka, et al.
Published: (2024)
by: Nayak, Nandeeka, et al.
Published: (2024)
Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models
by: Odemuyiwa, Toluwanimi O., et al.
Published: (2026)
by: Odemuyiwa, Toluwanimi O., et al.
Published: (2026)
Mapping Fusion: Improving FPGA Technology Mapping with ASIC Mapper
by: Yu, Cunxi
Published: (2025)
by: Yu, Cunxi
Published: (2025)
TeAAL: A Declarative Framework for Modeling Sparse Tensor Accelerators
by: Nayak, Nandeeka, et al.
Published: (2023)
by: Nayak, Nandeeka, et al.
Published: (2023)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
by: Tahmasebi, Faraz, et al.
Published: (2025)
by: Tahmasebi, Faraz, et al.
Published: (2025)
GDEV-AI: A Generalized Evaluation of Deep Learning Inference Scaling and Architectural Saturation
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
by: Rout, Nikhil, et al.
Published: (2025)
by: Rout, Nikhil, et al.
Published: (2025)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
by: Dudek, Piotr
Published: (2024)
by: Dudek, Piotr
Published: (2024)
Improving Injection-Throttling Mechanisms for Congestion Control for Data-center and Supercomputer Interconnects
by: Olmedilla, Cristina, et al.
Published: (2025)
by: Olmedilla, Cristina, et al.
Published: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
CLIPGen: A Chiplet Link IP Modeling and Generation Framework for 2.5D Architecture Exploration
by: Zhu, Zhengping, et al.
Published: (2026)
by: Zhu, Zhengping, et al.
Published: (2026)
Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
by: Li, Richie
Published: (2025)
by: Li, Richie
Published: (2025)
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
by: Özyılmaz, Mustafa Mert
Published: (2026)
by: Özyılmaz, Mustafa Mert
Published: (2026)
FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
by: Jia, Shuao, et al.
Published: (2025)
by: Jia, Shuao, et al.
Published: (2025)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness
by: Jiang, Zhe, et al.
Published: (2024)
by: Jiang, Zhe, et al.
Published: (2024)
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024)
by: Guan, Jiapeng, et al.
Published: (2024)
Lincoln AI Computing Survey (LAICS) and Trends
by: Reuther, Albert, et al.
Published: (2025)
by: Reuther, Albert, et al.
Published: (2025)
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
by: Li, Richie, et al.
Published: (2025)
by: Li, Richie, et al.
Published: (2025)
FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
by: Rahoof, Abdul, et al.
Published: (2025)
by: Rahoof, Abdul, et al.
Published: (2025)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
by: Xuan, Zihao, et al.
Published: (2026)
by: Xuan, Zihao, et al.
Published: (2026)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
SAT-MapIt: A SAT-based Modulo Scheduling Mapper for Coarse Grain Reconfigurable Architectures
by: Tirelli, Cristian, et al.
Published: (2025)
by: Tirelli, Cristian, et al.
Published: (2025)
Co-Design of Memory-Storage Systems for Workload Awareness with Interpretable Models
by: Sarkar, Jay, et al.
Published: (2026)
by: Sarkar, Jay, et al.
Published: (2026)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
A Compilation Framework for Quantum Circuits with Mid-Circuit Measurement Error Awareness
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
Enabling full-speed random access to the entire memory on the A100 GPU
by: Walker, Alden
Published: (2024)
by: Walker, Alden
Published: (2024)
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Fast and Practical Strassen's Matrix Multiplication using FPGAs
by: Ahmad, Afzal, et al.
Published: (2024)
by: Ahmad, Afzal, et al.
Published: (2024)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
by: Wang, Xuan, et al.
Published: (2024)
by: Wang, Xuan, et al.
Published: (2024)
Weight Block Sparsity: Training, Compilation, and AI Engine Accelerators
by: D'Alberto, Paolo, et al.
Published: (2024)
by: D'Alberto, Paolo, et al.
Published: (2024)
Accelerating Sensor Fusion in Neuromorphic Computing: A Case Study on Loihi-2
by: Isik, Murat, et al.
Published: (2024)
by: Isik, Murat, et al.
Published: (2024)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025)
by: Ta, Tuan, et al.
Published: (2025)
Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal
by: Jiang, Wenqi, et al.
Published: (2024)
by: Jiang, Wenqi, et al.
Published: (2024)
Similar Items
-
The Turbo-Charged Mapper: Fast and Optimal Mapping for Energy-efficient and Low-latency Accelerator Design
by: Gilbert, Michael, et al.
Published: (2026) -
CiMLoop: A Flexible, Accurate, and Fast Compute-In-Memory Modeling Tool
by: Andrulis, Tanner, et al.
Published: (2024) -
Modeling Analog-Digital-Converter Energy and Area for Compute-In-Memory Accelerator Design
by: Andrulis, Tanner, et al.
Published: (2024) -
Architecture-Level Modeling of Photonic Deep Neural Network Accelerators
by: Andrulis, Tanner, et al.
Published: (2024) -
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
by: Gilbert, Michael, et al.
Published: (2024)