Lifting to tensors when compiling scientific computing workloads for AI Engines
Fuente:
arXiv
Guardado en:
| Autores principales: | Brown, Nick, Rodriguez-Canal, Gabriel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Seamless acceleration of Fortran intrinsics via AMD AI engines
por: Brown, Nick, et al.
Publicado: (2025)
por: Brown, Nick, et al.
Publicado: (2025)
An MLIR pipeline for offloading Fortran to FPGAs via OpenMP
por: Rodriguez-Canal, Gabriel, et al.
Publicado: (2025)
por: Rodriguez-Canal, Gabriel, et al.
Publicado: (2025)
Investigations of multi-socket high core count RISC-V for HPC workloads
por: Brown, Nick, et al.
Publicado: (2025)
por: Brown, Nick, et al.
Publicado: (2025)
Cloud abstractions for AI workloads
por: Canini, Marco, et al.
Publicado: (2025)
por: Canini, Marco, et al.
Publicado: (2025)
Combining GPU and CPU for accelerating evolutionary computing workloads
por: Eynaliyev, Rustam, et al.
Publicado: (2025)
por: Eynaliyev, Rustam, et al.
Publicado: (2025)
Simulating LLM training workloads for heterogeneous compute and network infrastructure
por: Kumar, Sumit, et al.
Publicado: (2025)
por: Kumar, Sumit, et al.
Publicado: (2025)
Evaluating Versal AI Engines for option price discovery in market risk analysis
por: Klaisoongnoen, Mark, et al.
Publicado: (2024)
por: Klaisoongnoen, Mark, et al.
Publicado: (2024)
Fully integrating the Flang Fortran compiler with standard MLIR
por: Brown, Nick
Publicado: (2024)
por: Brown, Nick
Publicado: (2024)
A shared compilation stack for distributed-memory parallelism in stencil DSLs
por: Bisbas, George, et al.
Publicado: (2024)
por: Bisbas, George, et al.
Publicado: (2024)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
por: Merzky, Andre, et al.
Publicado: (2025)
por: Merzky, Andre, et al.
Publicado: (2025)
Is RISC-V ready for High Performance Computing? An evaluation of the Sophon SG2044
por: Brown, Nick
Publicado: (2025)
por: Brown, Nick
Publicado: (2025)
RISC-V for HPC: Where we are and where we need to go
por: Brown, Nick
Publicado: (2024)
por: Brown, Nick
Publicado: (2024)
RISC-V for HPC: An update of where we are and main action points
por: Brown, Nick
Publicado: (2025)
por: Brown, Nick
Publicado: (2025)
CRIU -- Checkpoint Restore in Userspace for computational simulations and scientific applications
por: Andrijauskas, Fabio, et al.
Publicado: (2024)
por: Andrijauskas, Fabio, et al.
Publicado: (2024)
Intelligent resource prediction for SAP HANA continuous integration build workloads
por: Mandel, Torsten, et al.
Publicado: (2026)
por: Mandel, Torsten, et al.
Publicado: (2026)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
por: Banchelli, Fabio, et al.
Publicado: (2025)
por: Banchelli, Fabio, et al.
Publicado: (2025)
Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator
por: Brown, Nick, et al.
Publicado: (2024)
por: Brown, Nick, et al.
Publicado: (2024)
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
por: Brown, Nick, et al.
Publicado: (2024)
por: Brown, Nick, et al.
Publicado: (2024)
Implementing OpenMP for Zig to enable its use in HPC context
por: Kacs, David, et al.
Publicado: (2024)
por: Kacs, David, et al.
Publicado: (2024)
Programming RISC-V accelerators via Fortran
por: Brown, Nick, et al.
Publicado: (2025)
por: Brown, Nick, et al.
Publicado: (2025)
Exploring Fast Fourier Transforms on the Tenstorrent Wormhole
por: Brown, Nick, et al.
Publicado: (2025)
por: Brown, Nick, et al.
Publicado: (2025)
Towards observability of scientific applications
por: Balis, Bartosz, et al.
Publicado: (2024)
por: Balis, Bartosz, et al.
Publicado: (2024)
Towards cloud-native scientific workflow management
por: Orzechowski, Michal, et al.
Publicado: (2024)
por: Orzechowski, Michal, et al.
Publicado: (2024)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
por: Zhao, Xinkui, et al.
Publicado: (2025)
por: Zhao, Xinkui, et al.
Publicado: (2025)
An approach to provide serverless scientific pipelines within the context of SKA
por: Ríos-Monje, Carlos, et al.
Publicado: (2023)
por: Ríos-Monje, Carlos, et al.
Publicado: (2023)
Batched DGEMMs for scientific codes running on long vector architectures
por: Banchelli, Fabio, et al.
Publicado: (2025)
por: Banchelli, Fabio, et al.
Publicado: (2025)
Evolving HPC services to enable ML workloads on HPE Cray EX
por: Schuppli, Stefano, et al.
Publicado: (2025)
por: Schuppli, Stefano, et al.
Publicado: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
por: Cheng, Ke, et al.
Publicado: (2024)
por: Cheng, Ke, et al.
Publicado: (2024)
SwarmSearch: Decentralized Search Engine with Self-Funding Economy
por: Gregoriadis, Marcel, et al.
Publicado: (2025)
por: Gregoriadis, Marcel, et al.
Publicado: (2025)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
por: Zhang, Zuoning, et al.
Publicado: (2024)
por: Zhang, Zuoning, et al.
Publicado: (2024)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
por: Biswas, Anish, et al.
Publicado: (2026)
por: Biswas, Anish, et al.
Publicado: (2026)
Optimizing the Longhorn Cloud-native Software Defined Storage Engine for High Performance
por: Kampadais, Konstantinos, et al.
Publicado: (2025)
por: Kampadais, Konstantinos, et al.
Publicado: (2025)
Comparing the Run-time Behavior of Modern PDES Engines on Alternative Hardware Architectures
por: Marotta, Romolo, et al.
Publicado: (2025)
por: Marotta, Romolo, et al.
Publicado: (2025)
Enabling Mixed criticality applications for the Versal AI-Engines
por: Sprave, Vincent, et al.
Publicado: (2026)
por: Sprave, Vincent, et al.
Publicado: (2026)
Developing a BLAS library for the AMD AI Engine
por: Laan, Tristan, et al.
Publicado: (2024)
por: Laan, Tristan, et al.
Publicado: (2024)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
por: Du, Kuntai, et al.
Publicado: (2025)
por: Du, Kuntai, et al.
Publicado: (2025)
Graph for Science: From API based Programming to Graph Engine based Programming for HPC
por: Zhang, Yu, et al.
Publicado: (2023)
por: Zhang, Yu, et al.
Publicado: (2023)
Multi-agent Reinforcement Learning-based In-place Scaling Engine for Edge-cloud Systems
por: Prodanov, Jovan, et al.
Publicado: (2025)
por: Prodanov, Jovan, et al.
Publicado: (2025)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
por: Qi, Shuyao, et al.
Publicado: (2026)
por: Qi, Shuyao, et al.
Publicado: (2026)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
por: Uchino, Yuki, et al.
Publicado: (2025)
por: Uchino, Yuki, et al.
Publicado: (2025)
Ejemplares similares
-
Seamless acceleration of Fortran intrinsics via AMD AI engines
por: Brown, Nick, et al.
Publicado: (2025) -
An MLIR pipeline for offloading Fortran to FPGAs via OpenMP
por: Rodriguez-Canal, Gabriel, et al.
Publicado: (2025) -
Investigations of multi-socket high core count RISC-V for HPC workloads
por: Brown, Nick, et al.
Publicado: (2025) -
Cloud abstractions for AI workloads
por: Canini, Marco, et al.
Publicado: (2025) -
Combining GPU and CPU for accelerating evolutionary computing workloads
por: Eynaliyev, Rustam, et al.
Publicado: (2025)