Kernel Looping: Eliminating Synchronization Boundaries for Peak Inference Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Koeplinger, David, Gandhi, Darshan, Nandkar, Pushkar, Sheeley, Nathan, Musaddiq, Matheen, Zhang, Leon, Goodbar, Reid, Shaffer, Matthew, Wang, Han, Wang, Angela, Wang, Mingran, Prabhakar, Raghu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
by: Prabhakar, Raghu, et al.
Published: (2024)
by: Prabhakar, Raghu, et al.
Published: (2024)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025)
by: Liu, Lian, et al.
Published: (2025)
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
by: Taji, Hossein, et al.
Published: (2025)
by: Taji, Hossein, et al.
Published: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
by: Li, Jonathan, et al.
Published: (2025)
by: Li, Jonathan, et al.
Published: (2025)
Synthesis-in-the-Loop Evaluation of LLMs for RTL Generation: Quality, Reliability, and Failure Modes
by: Fu, Weimin, et al.
Published: (2026)
by: Fu, Weimin, et al.
Published: (2026)
AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators
by: Agostini, Nicolas Bohm, et al.
Published: (2023)
by: Agostini, Nicolas Bohm, et al.
Published: (2023)
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024)
by: Guan, Jiapeng, et al.
Published: (2024)
DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing
by: Zhang, Yuhan, et al.
Published: (2026)
by: Zhang, Yuhan, et al.
Published: (2026)
Enabling full-speed random access to the entire memory on the A100 GPU
by: Walker, Alden
Published: (2024)
by: Walker, Alden
Published: (2024)
Fast and Practical Strassen's Matrix Multiplication using FPGAs
by: Ahmad, Afzal, et al.
Published: (2024)
by: Ahmad, Afzal, et al.
Published: (2024)
Veryl: A New Hardware Description Language as an Altarnative to SystemVerilog
by: Hatta, Naoya, et al.
Published: (2024)
by: Hatta, Naoya, et al.
Published: (2024)
Sequence-Based Incremental Concolic Testing of RTL Models
by: Witharana, Hasini, et al.
Published: (2023)
by: Witharana, Hasini, et al.
Published: (2023)
Non-interfering On-line and In-field SoC Testing
by: Strauch, Tobias
Published: (2024)
by: Strauch, Tobias
Published: (2024)
CIM-MLC: A Multi-level Compilation Stack for Computing-In-Memory Accelerators
by: Qu, Songyun, et al.
Published: (2024)
by: Qu, Songyun, et al.
Published: (2024)
AES-RV: Hardware-Efficient RISC-V Accelerator with Low-Latency AES Instruction Extension for IoT Security
by: Nguyen, Van Tinh, et al.
Published: (2025)
by: Nguyen, Van Tinh, et al.
Published: (2025)
HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
by: Suryadevara, Pranav
Published: (2025)
by: Suryadevara, Pranav
Published: (2025)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
by: Parameshwara, Arya, et al.
Published: (2025)
by: Parameshwara, Arya, et al.
Published: (2025)
ALL-MASK: A Reconfigurable Logic Locking Method for Multicore Architecture with Sequential-Instruction-Oriented Key
by: Wang, Jianfeng, et al.
Published: (2022)
by: Wang, Jianfeng, et al.
Published: (2022)
Megakernel vs Wavefront GPU Path Tracing
by: Padilla, Rafael, et al.
Published: (2026)
by: Padilla, Rafael, et al.
Published: (2026)
FPGA-based Acceleration for Convolutional Neural Networks: A Comprehensive Review
by: Jiang, Junye, et al.
Published: (2025)
by: Jiang, Junye, et al.
Published: (2025)
A Scalable Architecture for Efficient Multi-bit Fully Homomorphic Encryption
by: Ma, Jiaao, et al.
Published: (2025)
by: Ma, Jiaao, et al.
Published: (2025)
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
by: Lebold, Denis, et al.
Published: (2026)
by: Lebold, Denis, et al.
Published: (2026)
A Compilation Framework for Quantum Circuits with Mid-Circuit Measurement Error Awareness
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025)
by: Parameshwara, Arya
Published: (2025)
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories
by: Lee, Ming-Yen, et al.
Published: (2025)
by: Lee, Ming-Yen, et al.
Published: (2025)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
by: Su, Zhongling, et al.
Published: (2025)
by: Su, Zhongling, et al.
Published: (2025)
Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization
by: Ruggeri, Giuseppe, et al.
Published: (2025)
by: Ruggeri, Giuseppe, et al.
Published: (2025)
Revet: A Language and Compiler for Dataflow Threads
by: Rucker, Alexander, et al.
Published: (2023)
by: Rucker, Alexander, et al.
Published: (2023)
Optimization of 32-bit Unsigned Division by Constants on 64-bit Targets
by: Mitsunari, Shigeo, et al.
Published: (2026)
by: Mitsunari, Shigeo, et al.
Published: (2026)
EStacker: Explaining Battery-Less IoT System Performance with Energy Stacks
by: Liedtke, Lukas, et al.
Published: (2025)
by: Liedtke, Lukas, et al.
Published: (2025)
CLIPGen: A Chiplet Link IP Modeling and Generation Framework for 2.5D Architecture Exploration
by: Zhu, Zhengping, et al.
Published: (2026)
by: Zhu, Zhengping, et al.
Published: (2026)
Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness
by: Jiang, Zhe, et al.
Published: (2024)
by: Jiang, Zhe, et al.
Published: (2024)
Improved Prefetching Techniques for Linked Data Structures
by: Maruszewski, Nikola Vuk
Published: (2025)
by: Maruszewski, Nikola Vuk
Published: (2025)
FREESS: An Educational Simulator of a RISC-V-Inspired Superscalar Processor Based on Tomasulo's Algorithm
by: Giorgi, Roberto
Published: (2025)
by: Giorgi, Roberto
Published: (2025)
REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
by: Chen, Kangqi, et al.
Published: (2025)
by: Chen, Kangqi, et al.
Published: (2025)
Architectural Isolation as a Timing Safety Primitive for Edge AI Medical Devices: Controlled Experimental Evidence on a Shared-Silicon Platform
by: Swami, Akul Mallayya
Published: (2026)
by: Swami, Akul Mallayya
Published: (2026)
A Spatio-Temporal Graph Neural Networks Approach for Predicting Silent Data Corruption inducing Circuit-Level Faults
by: Wei, Shaoqi, et al.
Published: (2025)
by: Wei, Shaoqi, et al.
Published: (2025)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning
by: Das, Tamoghno, et al.
Published: (2025)
by: Das, Tamoghno, et al.
Published: (2025)
Similar Items
-
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
by: Prabhakar, Raghu, et al.
Published: (2024) -
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025) -
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
by: Taji, Hossein, et al.
Published: (2025) -
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
by: Li, Jonathan, et al.
Published: (2025) -
Synthesis-in-the-Loop Evaluation of LLMs for RTL Generation: Quality, Reliability, and Failure Modes
by: Fu, Weimin, et al.
Published: (2026)