Enabling mixed-precision with the help of tools: A Nekbone case study
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yanxiang, Castro, Pablo de Oliveira, Bientinesi, Paolo, Iakymchuk, Roman |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling mixed-precision in spectral element codes
by: Chen, Yanxiang, et al.
Published: (2025)
by: Chen, Yanxiang, et al.
Published: (2025)
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
by: Havdiak, Mykhailo, et al.
Published: (2024)
by: Havdiak, Mykhailo, et al.
Published: (2024)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024)
by: Li, Junjie
Published: (2024)
Network Centrality as a New Perspective on Microservice Architecture
by: Bakhtin, Alexander, et al.
Published: (2025)
by: Bakhtin, Alexander, et al.
Published: (2025)
gFaaS: Enabling Generic Functions in Serverless Computing
by: Chadha, Mohak, et al.
Published: (2024)
by: Chadha, Mohak, et al.
Published: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
by: Domke, Jens, et al.
Published: (2025)
by: Domke, Jens, et al.
Published: (2025)
Enabling MPI communication within Numba/LLVM JIT-compiled Python code using numba-mpi v1.0
by: Derlatka, Kacper, et al.
Published: (2024)
by: Derlatka, Kacper, et al.
Published: (2024)
ATOM: Asynchronous Training of Massive Models for Deep Learning in a Decentralized Environment
by: Wu, Xiaofeng, et al.
Published: (2024)
by: Wu, Xiaofeng, et al.
Published: (2024)
Metronome: Differentiated Delay Scheduling for Serverless Functions
by: Chen, Zhuangbin, et al.
Published: (2025)
by: Chen, Zhuangbin, et al.
Published: (2025)
TraceMesh: Scalable and Streaming Sampling for Distributed Traces
by: Chen, Zhuangbin, et al.
Published: (2024)
by: Chen, Zhuangbin, et al.
Published: (2024)
MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
Wherefore Art Thou? Provenance-Guided Automatic Online Debugging with Lumos
by: Chen, Jingyuan, et al.
Published: (2026)
by: Chen, Jingyuan, et al.
Published: (2026)
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
by: Yu, Guangba, et al.
Published: (2026)
by: Yu, Guangba, et al.
Published: (2026)
MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IR
by: Karchi, Jad El, et al.
Published: (2024)
by: Karchi, Jad El, et al.
Published: (2024)
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
by: Jiang, Zhihan, et al.
Published: (2025)
by: Jiang, Zhihan, et al.
Published: (2025)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
by: Zhu, Honglin, et al.
Published: (2026)
by: Zhu, Honglin, et al.
Published: (2026)
Supercharging Federated Learning with Flower and NVIDIA FLARE
by: Roth, Holger R., et al.
Published: (2024)
by: Roth, Holger R., et al.
Published: (2024)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
by: Singh, Abhinav, et al.
Published: (2023)
by: Singh, Abhinav, et al.
Published: (2023)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
High-Performance Star-M SVD for Big Data Compression
by: Hussain, Md Taufique, et al.
Published: (2026)
by: Hussain, Md Taufique, et al.
Published: (2026)
A shared compilation stack for distributed-memory parallelism in stencil DSLs
by: Bisbas, George, et al.
Published: (2024)
by: Bisbas, George, et al.
Published: (2024)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
by: Wang, Hansheng, et al.
Published: (2025)
by: Wang, Hansheng, et al.
Published: (2025)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)
by: Carrica, Vicki, et al.
Published: (2025)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
by: Bellavita, Julian, et al.
Published: (2026)
by: Bellavita, Julian, et al.
Published: (2026)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
by: Ham, David A., et al.
Published: (2024)
by: Ham, David A., et al.
Published: (2024)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
by: Machado, Rafael Ravedutti Lucio, et al.
Published: (2025)
by: Machado, Rafael Ravedutti Lucio, et al.
Published: (2025)
A new open source framework for multiscale modeling of fibrous materials on heterogeneous supercomputers
by: Merson, Jacob, et al.
Published: (2023)
by: Merson, Jacob, et al.
Published: (2023)
SYCL compute kernels for ExaHyPE
by: Loi, Chung Ming, et al.
Published: (2023)
by: Loi, Chung Ming, et al.
Published: (2023)
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
by: Schmid, Larissa, et al.
Published: (2024)
by: Schmid, Larissa, et al.
Published: (2024)
CloudHeatMap: Heatmap-Based Monitoring for Large-Scale Cloud Systems
by: Sohana, Sarah, et al.
Published: (2024)
by: Sohana, Sarah, et al.
Published: (2024)
$μ$OpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics
by: Japke, Nils, et al.
Published: (2025)
by: Japke, Nils, et al.
Published: (2025)
Adaptable TeaStore
by: Bliudze, Simon, et al.
Published: (2024)
by: Bliudze, Simon, et al.
Published: (2024)
A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPC
by: Sandås, Petter, et al.
Published: (2026)
by: Sandås, Petter, et al.
Published: (2026)
Building Castles in the Cloud: Architecting Resilient and Scalable Infrastructure
by: Gundla, Naresh Kumar
Published: (2024)
by: Gundla, Naresh Kumar
Published: (2024)
Histrio: a Serverless Actor System
by: Buttiglieri, Giorgio Natale, et al.
Published: (2024)
by: Buttiglieri, Giorgio Natale, et al.
Published: (2024)
Similar Items
-
Enabling mixed-precision in spectral element codes
by: Chen, Yanxiang, et al.
Published: (2025) -
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
by: Havdiak, Mykhailo, et al.
Published: (2024) -
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025) -
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024) -
Network Centrality as a New Perspective on Microservice Architecture
by: Bakhtin, Alexander, et al.
Published: (2025)