Mojo: MLIR-Based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Godoy, William F., Melnichenko, Tatiana, Valero-Lara, Pedro, Elwasif, Wael, Fackler, Philip, Da Silva, Rafael Ferreira, Teranishi, Keita, Vetter, Jeffrey S.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912604755066880
author Godoy, William F.
Melnichenko, Tatiana
Valero-Lara, Pedro
Elwasif, Wael
Fackler, Philip
Da Silva, Rafael Ferreira
Teranishi, Keita
Vetter, Jeffrey S.
author_facet Godoy, William F.
Melnichenko, Tatiana
Valero-Lara, Pedro
Elwasif, Wael
Fackler, Philip
Da Silva, Rafael Ferreira
Teranishi, Keita
Vetter, Jeffrey S.
contents We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM's Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python's interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree-Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo's performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mojo: MLIR-Based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem
Godoy, William F.
Melnichenko, Tatiana
Valero-Lara, Pedro
Elwasif, Wael
Fackler, Philip
Da Silva, Rafael Ferreira
Teranishi, Keita
Vetter, Jeffrey S.
Distributed, Parallel, and Cluster Computing
Computational Engineering, Finance, and Science
Emerging Technologies
Programming Languages
We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM's Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python's interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree-Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo's performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.
title Mojo: MLIR-Based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem
topic Distributed, Parallel, and Cluster Computing
Computational Engineering, Finance, and Science
Emerging Technologies
Programming Languages
url https://arxiv.org/abs/2509.21039