oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jianhui, Qin, Zhennan, Mei, Yijie, Cui, Jingze, Song, Yunfei, Chen, Ciyong, Zhang, Yifei, Du, Longsheng, Cheng, Xianhang, Jin, Baihui, Zhang, Yan, Ye, Jason, Lin, Eric, Lavery, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SENSEi: Input-Sensitive Compilation for Accelerating GNNs
by: Lenadora, Damitha, et al.
Published: (2023)
by: Lenadora, Damitha, et al.
Published: (2023)
Evaluating Compiler Optimization Impacts on zkVM Performance
by: Gassmann, Thomas, et al.
Published: (2025)
by: Gassmann, Thomas, et al.
Published: (2025)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
by: Zhang, Yijia, et al.
Published: (2024)
by: Zhang, Yijia, et al.
Published: (2024)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
by: You, Bozhi, et al.
Published: (2025)
by: You, Bozhi, et al.
Published: (2025)
Priority Sampling of Large Language Models for Compilers
by: Grubisic, Dejan, et al.
Published: (2024)
by: Grubisic, Dejan, et al.
Published: (2024)
The Next 700 ML-Enabled Compiler Optimizations
by: VenkataKeerthy, S., et al.
Published: (2023)
by: VenkataKeerthy, S., et al.
Published: (2023)
On Combining Two Server Control Policies for Energy Efficiency
by: Dai, Jingze, et al.
Published: (2025)
by: Dai, Jingze, et al.
Published: (2025)
An Empirical Study on the Performance and Energy Usage of Compiled Python Code
by: Stoico, Vincenzo, et al.
Published: (2025)
by: Stoico, Vincenzo, et al.
Published: (2025)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
by: Sakib, Nazmus, et al.
Published: (2025)
by: Sakib, Nazmus, et al.
Published: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
by: Szafarczyk, Robert, et al.
Published: (2025)
by: Szafarczyk, Robert, et al.
Published: (2025)
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
by: Pan, Haolin, et al.
Published: (2026)
by: Pan, Haolin, et al.
Published: (2026)
Versatile Cross-platform Compilation Toolchain for Schrödinger-style Quantum Circuit Simulation
by: Lu, Yuncheng, et al.
Published: (2025)
by: Lu, Yuncheng, et al.
Published: (2025)
LightningSimV2: Faster and Scalable Simulation for High-Level Synthesis via Graph Compilation and Optimization
by: Sarkar, Rishov, et al.
Published: (2024)
by: Sarkar, Rishov, et al.
Published: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
by: Gupta, Ahan, et al.
Published: (2026)
by: Gupta, Ahan, et al.
Published: (2026)
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
QOPS: A Compiler Framework for Quantum Circuit Simulation Acceleration with Profile Guided Optimizations
by: Wu, Yu-Tsung, et al.
Published: (2024)
by: Wu, Yu-Tsung, et al.
Published: (2024)
VecTrans: Enhancing Compiler Auto-Vectorization through LLM-Assisted Code Transformations
by: Zheng, Zhongchun, et al.
Published: (2025)
by: Zheng, Zhongchun, et al.
Published: (2025)
ACPO: AI-Enabled Compiler Framework
by: Ashouri, Amir H., et al.
Published: (2023)
by: Ashouri, Amir H., et al.
Published: (2023)
Library Liberation: Competitive Performance Matmul Through Compiler-composed Nanokernels
by: Thangamani, Arun, et al.
Published: (2025)
by: Thangamani, Arun, et al.
Published: (2025)
KeyMemRT Compiler and Runtime: Unlocking Memory-Scalable FHE
by: Ünay, Eymen, et al.
Published: (2026)
by: Ünay, Eymen, et al.
Published: (2026)
AFarePart: Accuracy-aware Fault-resilient Partitioner for DNN Edge Accelerators
by: Debnath, Mukta, et al.
Published: (2025)
by: Debnath, Mukta, et al.
Published: (2025)
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
by: Le, Truong-Thanh, et al.
Published: (2026)
by: Le, Truong-Thanh, et al.
Published: (2026)
Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models
by: Rosas, Miguel Romero, et al.
Published: (2024)
by: Rosas, Miguel Romero, et al.
Published: (2024)
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
by: Yang, Peiming, et al.
Published: (2025)
by: Yang, Peiming, et al.
Published: (2025)
Developing a Modular Compiler for a Subset of a C-like Language
by: Dutta, Debasish, et al.
Published: (2025)
by: Dutta, Debasish, et al.
Published: (2025)
Compiler-First State Space Duality and Portable $O(1)$ Autoregressive Caching for Inference
by: Santoni, Cosmo
Published: (2026)
by: Santoni, Cosmo
Published: (2026)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)
by: Fu, Zizhuo, et al.
Published: (2025)
CarbonCP: Carbon-Aware DNN Partitioning with Conformal Prediction for Sustainable Edge Intelligence
by: Ke, Hongyu, et al.
Published: (2024)
by: Ke, Hongyu, et al.
Published: (2024)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
by: Huang, Haochen, et al.
Published: (2025)
by: Huang, Haochen, et al.
Published: (2025)
Multi-DNN Inference of Sparse Models on Edge SoCs
by: Luo, Jiawei, et al.
Published: (2026)
by: Luo, Jiawei, et al.
Published: (2026)
Integration of a systolic array based hardware accelerator into a DNN operator auto-tuning framework
by: Peccia, F. N., et al.
Published: (2022)
by: Peccia, F. N., et al.
Published: (2022)
Protean Compiler: An Agile Framework to Drive Fine-grain Phase Ordering
by: Ashouri, Amir H., et al.
Published: (2026)
by: Ashouri, Amir H., et al.
Published: (2026)
Efficient Hybrid Amplitude-Phase Quantization for Multi-Antenna Relay System
by: Kim, Changdae, et al.
Published: (2025)
by: Kim, Changdae, et al.
Published: (2025)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
by: Zhang, Niansong, et al.
Published: (2025)
by: Zhang, Niansong, et al.
Published: (2025)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026)
by: Ding, Jiabiao, et al.
Published: (2026)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
by: Dagli, Ismet, et al.
Published: (2023)
by: Dagli, Ismet, et al.
Published: (2023)
Optimizing Winograd Convolution on ARMv8 processors
by: Gui, Haoyuan, et al.
Published: (2024)
by: Gui, Haoyuan, et al.
Published: (2024)
Reducing Waiting Time for Medical Tourists Through Hybrid Agent-Based and Discrete-Event Simulation: A Hospital Case Study
by: Baghi, Melika, et al.
Published: (2026)
by: Baghi, Melika, et al.
Published: (2026)
Accurate and Scalable Many-Node Simulation
by: Eyerman, Stijn, et al.
Published: (2024)
by: Eyerman, Stijn, et al.
Published: (2024)
Similar Items
-
SENSEi: Input-Sensitive Compilation for Accelerating GNNs
by: Lenadora, Damitha, et al.
Published: (2023) -
Evaluating Compiler Optimization Impacts on zkVM Performance
by: Gassmann, Thomas, et al.
Published: (2025) -
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
by: Zhang, Yijia, et al.
Published: (2024) -
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
by: You, Bozhi, et al.
Published: (2025) -
Priority Sampling of Large Language Models for Compilers
by: Grubisic, Dejan, et al.
Published: (2024)