Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Georganas, Evangelos, Kalamkar, Dhiraj, Voronin, Kirill, Kundu, Abhisek, Noack, Antonio, Pabst, Hans, Breuer, Alexander, Heinecke, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple
by: Georganas, Evangelos, et al.
Published: (2026)
by: Georganas, Evangelos, et al.
Published: (2026)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
by: Remke, Stefan, et al.
Published: (2024)
by: Remke, Stefan, et al.
Published: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Towards a high-performance AI compiler with upstream MLIR
by: Golin, Renato, et al.
Published: (2024)
by: Golin, Renato, et al.
Published: (2024)
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
by: Brown, Nick, et al.
Published: (2024)
by: Brown, Nick, et al.
Published: (2024)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
by: Bridges, Patrick G., et al.
Published: (2026)
by: Bridges, Patrick G., et al.
Published: (2026)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
by: Tharwani, Jay, et al.
Published: (2025)
by: Tharwani, Jay, et al.
Published: (2025)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
by: Zhang, Jinxiao, et al.
Published: (2026)
by: Zhang, Jinxiao, et al.
Published: (2026)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
by: Chazapis, Antony, et al.
Published: (2024)
by: Chazapis, Antony, et al.
Published: (2024)
An Autonomy Loop for Dynamic HPC Job Time Limit Adjustment
by: Jakobsche, Thomas, et al.
Published: (2025)
by: Jakobsche, Thomas, et al.
Published: (2025)
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
by: Vooturi, Dharma Teja, et al.
Published: (2026)
by: Vooturi, Dharma Teja, et al.
Published: (2026)
Autonomy Loops for Monitoring, Operational Data Analytics, Feedback, and Response in HPC Operations
by: Boito, Francieli, et al.
Published: (2024)
by: Boito, Francieli, et al.
Published: (2024)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
by: Yi, Xinyao
Published: (2024)
by: Yi, Xinyao
Published: (2024)
Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure
by: Sfiligoi, Igor, et al.
Published: (2025)
by: Sfiligoi, Igor, et al.
Published: (2025)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
by: Miyashita, Yusuke, et al.
Published: (2024)
by: Miyashita, Yusuke, et al.
Published: (2024)
DOLMA: A Data Object Level Memory Disaggregation Framework for HPC Applications
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
An Incremental Multi-Level, Multi-Scale Approach to Assessment of Multifidelity HPC Systems
by: Shilpika, Shilpika, et al.
Published: (2025)
by: Shilpika, Shilpika, et al.
Published: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
by: Barrak, Amine, et al.
Published: (2025)
by: Barrak, Amine, et al.
Published: (2025)
Driving Computational Efficiency in Large-Scale Platforms using HPC Technologies
by: Mendez, Alexander Martinez, et al.
Published: (2026)
by: Mendez, Alexander Martinez, et al.
Published: (2026)
An Analysis of HPC and Edge Architectures in the Cloud
by: Santillan, Steven, et al.
Published: (2025)
by: Santillan, Steven, et al.
Published: (2025)
Towards CXL Resilience to CPU Failures
by: Psistakis, Antonis, et al.
Published: (2026)
by: Psistakis, Antonis, et al.
Published: (2026)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
by: Spoczynski, Marcin, et al.
Published: (2026)
by: Spoczynski, Marcin, et al.
Published: (2026)
HPC with Enhanced User Separation
by: Prout, Andrew, et al.
Published: (2024)
by: Prout, Andrew, et al.
Published: (2024)
Analysis of the carbon footprint of HPC
by: Benhari, Abdessalam, et al.
Published: (2025)
by: Benhari, Abdessalam, et al.
Published: (2025)
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
by: Arima, Eishi, et al.
Published: (2024)
by: Arima, Eishi, et al.
Published: (2024)
WindVE: Collaborative CPU-NPU Vector Embedding
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Combining GPU and CPU for accelerating evolutionary computing workloads
by: Eynaliyev, Rustam, et al.
Published: (2025)
by: Eynaliyev, Rustam, et al.
Published: (2025)
MRSch: Multi-Resource Scheduling for HPC
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
HPC-vQPU: A Service-Export Architecture for Virtual QPUs on Batch-Scheduled HPC Systems
by: Liu, Shusen, et al.
Published: (2026)
by: Liu, Shusen, et al.
Published: (2026)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
by: Zhang, Lingqi, et al.
Published: (2025)
by: Zhang, Lingqi, et al.
Published: (2025)
Architectural Foundations for Checkpointing and Restoration in Quantum HPC Systems
by: Guan, Qiang, et al.
Published: (2026)
by: Guan, Qiang, et al.
Published: (2026)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
by: Ding, Zhimin, et al.
Published: (2024)
by: Ding, Zhimin, et al.
Published: (2024)
UNR: Unified Notifiable RMA Library for HPC
by: Feng, Guangnan, et al.
Published: (2024)
by: Feng, Guangnan, et al.
Published: (2024)
An Elastic Job Scheduler for HPC Applications on the Cloud
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Sarus Suite: Cloud-native Containers for HPC
by: Madonna, Alberto, et al.
Published: (2026)
by: Madonna, Alberto, et al.
Published: (2026)
Wilkins: HPC In Situ Workflows Made Easy
by: Yildiz, Orcun, et al.
Published: (2024)
by: Yildiz, Orcun, et al.
Published: (2024)
Similar Items
-
Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple
by: Georganas, Evangelos, et al.
Published: (2026) -
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
by: Remke, Stefan, et al.
Published: (2024) -
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024) -
Towards a high-performance AI compiler with upstream MLIR
by: Golin, Renato, et al.
Published: (2024) -
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
by: Brown, Nick, et al.
Published: (2024)