Compiler-supported reduced precision and AoS-SoA transformations for heterogeneous hardware
Fuente:
arXiv
Saved in:
| Main Authors: | Radtke, Pawel K., Weinzierl, Tobias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Annotation-guided AoS-to-SoA conversions and GPU offloading with data views in C++
by: Radtke, Pawel K., et al.
Published: (2025)
by: Radtke, Pawel K., et al.
Published: (2025)
Compiler support for semi-manual AoS-to-SoA conversions with data views
by: Radtke, Pawel K., et al.
Published: (2024)
by: Radtke, Pawel K., et al.
Published: (2024)
SYCL compute kernels for ExaHyPE
by: Loi, Chung Ming, et al.
Published: (2023)
by: Loi, Chung Ming, et al.
Published: (2023)
GPU Implementations for Midsize Integer Addition and Multiplication
by: Oancea, Cosmin E., et al.
Published: (2024)
by: Oancea, Cosmin E., et al.
Published: (2024)
Julia GraphBLAS with Nonblocking Execution
by: Costanza, Pascal, et al.
Published: (2025)
by: Costanza, Pascal, et al.
Published: (2025)
Annotation‐Guided AoS‐to‐SoA Conversions and GPU Offloading With Data Views in C++
by: Pawel K. Radtke, et al.
Published: (2025)
by: Pawel K. Radtke, et al.
Published: (2025)
Enabling mixed-precision in spectral element codes
by: Chen, Yanxiang, et al.
Published: (2025)
by: Chen, Yanxiang, et al.
Published: (2025)
Detrimental task execution patterns in mainstream OpenMP runtimes
by: Tuft, Adam S., et al.
Published: (2024)
by: Tuft, Adam S., et al.
Published: (2024)
A new open source framework for multiscale modeling of fibrous materials on heterogeneous supercomputers
by: Merson, Jacob, et al.
Published: (2023)
by: Merson, Jacob, et al.
Published: (2023)
Enabling mixed-precision with the help of tools: A Nekbone case study
by: Chen, Yanxiang, et al.
Published: (2024)
by: Chen, Yanxiang, et al.
Published: (2024)
RVISmith: Fuzzing Compilers for RVV Intrinsics
by: He, Yibo, et al.
Published: (2025)
by: He, Yibo, et al.
Published: (2025)
A shared compilation stack for distributed-memory parallelism in stencil DSLs
by: Bisbas, George, et al.
Published: (2024)
by: Bisbas, George, et al.
Published: (2024)
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
by: Havdiak, Mykhailo, et al.
Published: (2024)
by: Havdiak, Mykhailo, et al.
Published: (2024)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
by: Singh, Abhinav, et al.
Published: (2023)
by: Singh, Abhinav, et al.
Published: (2023)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
High-Performance Star-M SVD for Big Data Compression
by: Hussain, Md Taufique, et al.
Published: (2026)
by: Hussain, Md Taufique, et al.
Published: (2026)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
by: Zhu, Honglin, et al.
Published: (2026)
by: Zhu, Honglin, et al.
Published: (2026)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
by: Wang, Hansheng, et al.
Published: (2025)
by: Wang, Hansheng, et al.
Published: (2025)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)
by: Carrica, Vicki, et al.
Published: (2025)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
by: Bellavita, Julian, et al.
Published: (2026)
by: Bellavita, Julian, et al.
Published: (2026)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
by: Ham, David A., et al.
Published: (2024)
by: Ham, David A., et al.
Published: (2024)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
by: Machado, Rafael Ravedutti Lucio, et al.
Published: (2025)
by: Machado, Rafael Ravedutti Lucio, et al.
Published: (2025)
Enabling MPI communication within Numba/LLVM JIT-compiled Python code using numba-mpi v1.0
by: Derlatka, Kacper, et al.
Published: (2024)
by: Derlatka, Kacper, et al.
Published: (2024)
Fray: An Efficient General-Purpose Concurrency Testing Platform for the JVM (Extended Version)
by: Li, Ao, et al.
Published: (2025)
by: Li, Ao, et al.
Published: (2025)
Determinacy with Priorities up to Clocks
by: Liquori, Luigi, et al.
Published: (2026)
by: Liquori, Luigi, et al.
Published: (2026)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Comparative analysis of large data processing in Apache Spark using Java, Python and Scala
by: Borodii, Ivan, et al.
Published: (2025)
by: Borodii, Ivan, et al.
Published: (2025)
Automated MPI-X code generation for scalable finite-difference solvers
by: Bisbas, George, et al.
Published: (2023)
by: Bisbas, George, et al.
Published: (2023)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
by: Woller, Fabian, et al.
Published: (2025)
by: Woller, Fabian, et al.
Published: (2025)
Performance measurements of modern Fortran MPI applications with Score-P
by: Corbin, Gregor
Published: (2025)
by: Corbin, Gregor
Published: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
by: Panova, Elena, et al.
Published: (2022)
by: Panova, Elena, et al.
Published: (2022)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
by: Tu, Jiqun, et al.
Published: (2026)
by: Tu, Jiqun, et al.
Published: (2026)
ScanWeaver: Compiler-Driven Parallelization of Affine Recurrences via Associative Scan Lowering
by: Wu, Qiying, et al.
Published: (2026)
by: Wu, Qiying, et al.
Published: (2026)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
Similar Items
-
Annotation-guided AoS-to-SoA conversions and GPU offloading with data views in C++
by: Radtke, Pawel K., et al.
Published: (2025) -
Compiler support for semi-manual AoS-to-SoA conversions with data views
by: Radtke, Pawel K., et al.
Published: (2024) -
SYCL compute kernels for ExaHyPE
by: Loi, Chung Ming, et al.
Published: (2023) -
GPU Implementations for Midsize Integer Addition and Multiplication
by: Oancea, Cosmin E., et al.
Published: (2024) -
Julia GraphBLAS with Nonblocking Execution
by: Costanza, Pascal, et al.
Published: (2025)