Performance Portable Gradient Computations Using Source Transformation
Fuente:
arXiv
Saved in:
| Main Authors: | Liegeois, Kim, Kelley, Brian, Phipps, Eric, Rajamanickam, Sivasankaran, Vassilev, Vassil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Performance Portable Matrix Free Dense MTTKRP in GenTen
by: Kosmacher, Gabriel, et al.
Published: (2025)
by: Kosmacher, Gabriel, et al.
Published: (2025)
LAPIS: A Performance Portable, High Productivity Compiler Framework
by: Kelley, Brian, et al.
Published: (2025)
by: Kelley, Brian, et al.
Published: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
A Performance-Portable, Massively Parallel Distributed Nonuniform FFT
by: Fischill, Paul, et al.
Published: (2026)
by: Fischill, Paul, et al.
Published: (2026)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)
by: Carrica, Vicki, et al.
Published: (2025)
Transformations of Computational Meshes
by: Knepley, Matthew G.
Published: (2025)
by: Knepley, Matthew G.
Published: (2025)
Computing FFTs at Target Precision Using Lower-Precision FFTs
by: Kawakami, Shota, et al.
Published: (2026)
by: Kawakami, Shota, et al.
Published: (2026)
GeoFlood (v1.0.0): Computational model for overland flooding
by: Kyanjo, Brian, et al.
Published: (2024)
by: Kyanjo, Brian, et al.
Published: (2024)
Towards Richer Challenge Problems for Scientific Computing Correctness
by: Sottile, Matthew, et al.
Published: (2025)
by: Sottile, Matthew, et al.
Published: (2025)
Trilinos: Enabling Scientific Computing Across Diverse Hardware Architectures at Scale
by: Mayr, Matthias, et al.
Published: (2025)
by: Mayr, Matthias, et al.
Published: (2025)
Hiperwalk: Simulation of Quantum Walks with Heterogeneous High-Performance Computing
by: Motta, Paulo, et al.
Published: (2024)
by: Motta, Paulo, et al.
Published: (2024)
f4ncgb: High Performance Gröbner Basis Computations in Free Algebras
by: Heisinger, Maximilian, et al.
Published: (2025)
by: Heisinger, Maximilian, et al.
Published: (2025)
Performant Tridiagonal Factorization of Skew-Symmetric Matrices
by: Satyarth, Ishna, et al.
Published: (2024)
by: Satyarth, Ishna, et al.
Published: (2024)
Analyzing Computational Approaches for Differential Equations: A Study of MATLAB, Mathematica, and Maple
by: Ogethakpo, Arhonefe Joseph, et al.
Published: (2025)
by: Ogethakpo, Arhonefe Joseph, et al.
Published: (2025)
A Novel SIMD-Optimized Implementation for Fast and Memory-Efficient Trigonometric Computation
by: Goyal, Nikhil Dev, et al.
Published: (2025)
by: Goyal, Nikhil Dev, et al.
Published: (2025)
Efficient Computation of Collatz Sequence Stopping Times: A Novel Algorithmic Approach
by: Getachew, Eyob Solomon, et al.
Published: (2025)
by: Getachew, Eyob Solomon, et al.
Published: (2025)
Open Source Prover in the Attic
by: Kovács, Zoltán, et al.
Published: (2024)
by: Kovács, Zoltán, et al.
Published: (2024)
Correctly Rounded Functions For Vector Applications: A Performance Study
by: Anderson, Cristina, et al.
Published: (2026)
by: Anderson, Cristina, et al.
Published: (2026)
Performance Analysis of Effective Symbolic Methods for Solving Band Matrix SLAEs
by: Veneva, Milena, et al.
Published: (2019)
by: Veneva, Milena, et al.
Published: (2019)
svds-C: A Multi-Thread C Code for Computing Truncated Singular Value Decomposition
by: Feng, Xu, et al.
Published: (2024)
by: Feng, Xu, et al.
Published: (2024)
RLibm-MultiRound: Correctly Rounded Math Libraries Without Worrying about the Application's Rounding Mode
by: Park, Sehyeok, et al.
Published: (2025)
by: Park, Sehyeok, et al.
Published: (2025)
Minimization of Nonlinear Energies in Python Using FEM and Automatic Differentiation Tools
by: Béreš, Michal, et al.
Published: (2024)
by: Béreš, Michal, et al.
Published: (2024)
Odd but Error-Free FastTwoSum: More General Conditions for FastTwoSum as an Error-Free Transformation for Faithful Rounding Modes
by: Park, Sehyeok, et al.
Published: (2026)
by: Park, Sehyeok, et al.
Published: (2026)
Fast Large-Scale Model-Based Iterative Tomography via Exploiting Mathematical Structure, Hierarchical Optimization, Smart Initialization, and Distributed GPU Computing
by: Kumar, Dinesh, et al.
Published: (2026)
by: Kumar, Dinesh, et al.
Published: (2026)
Sparse Iterative Solvers Using High-Precision Arithmetic with Quasi Multi-Word Algorithms
by: Mukunoki, Daichi, et al.
Published: (2025)
by: Mukunoki, Daichi, et al.
Published: (2025)
Interface for Sparse Linear Algebra Operations
by: Abdelfattah, Ahmad, et al.
Published: (2024)
by: Abdelfattah, Ahmad, et al.
Published: (2024)
Sparse Automatic Differentiation for Complex Networks of Differential-Algebraic Equations Using Abstract Elementary Algebra
by: Peles, Slaven, et al.
Published: (2015)
by: Peles, Slaven, et al.
Published: (2015)
A Computational Framework and Implementation of Implicit Priors in Bayesian Inverse Problems
by: Everink, Jasper M., et al.
Published: (2025)
by: Everink, Jasper M., et al.
Published: (2025)
Improving Runtime Performance of Tensor Computations using Rust From Python
by: Harding, Kimmie, et al.
Published: (2025)
by: Harding, Kimmie, et al.
Published: (2025)
Dimensional Peeking for Low-Variance Gradients in Zeroth-Order Discrete Optimization via Simulation
by: Andelfinger, Philipp, et al.
Published: (2026)
by: Andelfinger, Philipp, et al.
Published: (2026)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
by: Scheffler, Paul, et al.
Published: (2024)
by: Scheffler, Paul, et al.
Published: (2024)
Integration of Quantum Accelerators with High Performance Computing -- A Review of Quantum Programming Tools
by: Elsharkawy, Amr, et al.
Published: (2023)
by: Elsharkawy, Amr, et al.
Published: (2023)
InfoFusion Controller: Informed TRRT Star with Mutual Information based on Fusion of Pure Pursuit and MPC for Enhanced Path Planning
by: Choi, Seongjun, et al.
Published: (2025)
by: Choi, Seongjun, et al.
Published: (2025)
QCLAB: A Matlab Toolbox for Quantum Computing
by: Keip, Sophia, et al.
Published: (2025)
by: Keip, Sophia, et al.
Published: (2025)
Faster Base64 Encoding and Decoding Using AVX2 Instructions
by: Muła, Wojciech, et al.
Published: (2017)
by: Muła, Wojciech, et al.
Published: (2017)
Experience converting a large mathematical software package written in C++ to C++20 modules
by: Bangerth, Wolfgang
Published: (2025)
by: Bangerth, Wolfgang
Published: (2025)
Enhancing non-Perl bioinformatic applications with Perl: Building novel, component based applications using Object Orientation, PDL, Alien, FFI, Inline and OpenMP
by: Argyropoulos, Christos
Published: (2024)
by: Argyropoulos, Christos
Published: (2024)
Portability of Fortran's `do concurrent' on GPUs
by: Caplan, Ronald M., et al.
Published: (2024)
by: Caplan, Ronald M., et al.
Published: (2024)
Jet: Multilevel Graph Partitioning on Graphics Processing Units
by: Gilbert, Michael S., et al.
Published: (2023)
by: Gilbert, Michael S., et al.
Published: (2023)
Similar Items
-
A Performance Portable Matrix Free Dense MTTKRP in GenTen
by: Kosmacher, Gabriel, et al.
Published: (2025) -
LAPIS: A Performance Portable, High Productivity Compiler Framework
by: Kelley, Brian, et al.
Published: (2025) -
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025) -
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025) -
A Performance-Portable, Massively Parallel Distributed Nonuniform FFT
by: Fischill, Paul, et al.
Published: (2026)