Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
Fuente:
arXiv
Saved in:
| Main Authors: | Kulkarni, Kaushik, Klöckner, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Canonicalization of Batched Einstein Summations for Tuning Retrieval
by: Kulkarni, Kaushik, et al.
Published: (2026)
by: Kulkarni, Kaushik, et al.
Published: (2026)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
by: Tu, Jiqun, et al.
Published: (2026)
by: Tu, Jiqun, et al.
Published: (2026)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
by: Panova, Elena, et al.
Published: (2022)
by: Panova, Elena, et al.
Published: (2022)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Automated MPI-X code generation for scalable finite-difference solvers
by: Bisbas, George, et al.
Published: (2023)
by: Bisbas, George, et al.
Published: (2023)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
by: Woller, Fabian, et al.
Published: (2025)
by: Woller, Fabian, et al.
Published: (2025)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Performance measurements of modern Fortran MPI applications with Score-P
by: Corbin, Gregor
Published: (2025)
by: Corbin, Gregor
Published: (2025)
Scaling the memory wall using mixed-precision -- HPG-MxP on an exascale machine
by: Kashi, Aditya, et al.
Published: (2025)
by: Kashi, Aditya, et al.
Published: (2025)
Adaptive time step selection for Spectral Deferred Correction
by: Saupe, Thomas, et al.
Published: (2024)
by: Saupe, Thomas, et al.
Published: (2024)
Resilience Against Soft Faults through Adaptivity in Spectral Deferred Correction
by: Saupe, Thomas, et al.
Published: (2024)
by: Saupe, Thomas, et al.
Published: (2024)
pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
by: Tuteja, Keshvi, et al.
Published: (2025)
by: Tuteja, Keshvi, et al.
Published: (2025)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
by: Li, Junjie
Published: (2024)
by: Li, Junjie
Published: (2024)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)
by: Kriemann, Ronald
Published: (2024)
Exploiting nested task-parallelism in the $\mathcal{H}-LU$ factorization
by: Carratalá-Sáez, Rocío, et al.
Published: (2019)
by: Carratalá-Sáez, Rocío, et al.
Published: (2019)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
by: Owen, Herbert, et al.
Published: (2024)
by: Owen, Herbert, et al.
Published: (2024)
Ridgeline: A 2D Roofline Model for Distributed Systems
by: Checconi, Fabio, et al.
Published: (2022)
by: Checconi, Fabio, et al.
Published: (2022)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
by: Carrica, Vicki, et al.
Published: (2026)
by: Carrica, Vicki, et al.
Published: (2026)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
by: Ham, David A., et al.
Published: (2024)
by: Ham, David A., et al.
Published: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024)
by: Afzal, Ayesha, et al.
Published: (2024)
Parallel Sparse and Data-Sparse Factorization-based Linear Solvers
by: Li, Xiaoye Sherry, et al.
Published: (2026)
by: Li, Xiaoye Sherry, et al.
Published: (2026)
A multigrid reduction framework for domains with symmetries
by: Alsalti-Baldellou, Àdel, et al.
Published: (2024)
by: Alsalti-Baldellou, Àdel, et al.
Published: (2024)
Fully-Automated Code Generation for Efficient Computation of Sparse Matrix Permanents on GPUs
by: Elbek, Deniz, et al.
Published: (2025)
by: Elbek, Deniz, et al.
Published: (2025)
Parallel Gauss-Jordan Elimination and System Reduction for Efficient Circuit Simulation
by: Noveski, Filip, et al.
Published: (2026)
by: Noveski, Filip, et al.
Published: (2026)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
by: Venkat, Sreeram, et al.
Published: (2025)
by: Venkat, Sreeram, et al.
Published: (2025)
Communication-Efficient and Memory-Aware Parallel Bootstrapping using MPI
by: Zhang, Di
Published: (2025)
by: Zhang, Di
Published: (2025)
How to Rent GPUs on a Budget
by: Li, Zhouzi, et al.
Published: (2024)
by: Li, Zhouzi, et al.
Published: (2024)
Parallel performance of shared memory parallel spectral deferred corrections
by: Freese, Philip, et al.
Published: (2024)
by: Freese, Philip, et al.
Published: (2024)
Easy Acceleration with Distributed Arrays
by: Kepner, Jeremy, et al.
Published: (2025)
by: Kepner, Jeremy, et al.
Published: (2025)
Near Real-time Adaptive Isotropic and Anisotropic Image-to-mesh Conversion for Numerical Simulations Involving Cerebral Aneurysms
by: Garner, Kevin, et al.
Published: (2024)
by: Garner, Kevin, et al.
Published: (2024)
ML-Based Optimum Number of CUDA Streams for the GPU Implementation of the Tridiagonal Partition Method
by: Veneva, Milena, et al.
Published: (2025)
by: Veneva, Milena, et al.
Published: (2025)
ML-Based Optimum Sub-system Size Heuristic for the GPU Implementation of the Tridiagonal Partition Method
by: Veneva, Milena
Published: (2025)
by: Veneva, Milena
Published: (2025)
Nearest Neighbors GParareal: Improving Scalability of Gaussian Processes for Parallel-in-Time Solvers
by: Gattiglio, Guglielmo, et al.
Published: (2024)
by: Gattiglio, Guglielmo, et al.
Published: (2024)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
High-level Stream Processing: A Complementary Analysis of Fault Recovery
by: Vogel, Adriano, et al.
Published: (2024)
by: Vogel, Adriano, et al.
Published: (2024)
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2024)
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2024)
Similar Items
-
Canonicalization of Batched Einstein Summations for Tuning Retrieval
by: Kulkarni, Kaushik, et al.
Published: (2026) -
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
by: Tu, Jiqun, et al.
Published: (2026) -
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
by: Panova, Elena, et al.
Published: (2022) -
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
by: Katagiri, Takahiro, et al.
Published: (2024) -
Automated MPI-X code generation for scalable finite-difference solvers
by: Bisbas, George, et al.
Published: (2023)