Canonicalization of Batched Einstein Summations for Tuning Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Kulkarni, Kaushik, Klöckner, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
by: Kulkarni, Kaushik, et al.
Published: (2025)
by: Kulkarni, Kaushik, et al.
Published: (2025)
Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
by: Martinez-Ferrer, Pedro J., et al.
Published: (2025)
by: Martinez-Ferrer, Pedro J., et al.
Published: (2025)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Bandicoot: A Templated C++ Library for GPU Linear Algebra
by: Curtin, Ryan R., et al.
Published: (2025)
by: Curtin, Ryan R., et al.
Published: (2025)
Exceeding the Numerical and Performance Characteristics of IEEE-754 SGEMM with BFloat16 Tensor Cores on GPUs for Scientific Computing
by: Bayraktar, Harun, et al.
Published: (2026)
by: Bayraktar, Harun, et al.
Published: (2026)
Fast GPU Linear Algebra via Compile Time Expression Fusion
by: Curtin, Ryan R., et al.
Published: (2026)
by: Curtin, Ryan R., et al.
Published: (2026)
Armadillo: An Efficient Framework for Numerical Linear Algebra
by: Sanderson, Conrad, et al.
Published: (2025)
by: Sanderson, Conrad, et al.
Published: (2025)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
by: Chillarón, Mónica, et al.
Published: (2024)
by: Chillarón, Mónica, et al.
Published: (2024)
Trilinos: Enabling Scientific Computing Across Diverse Hardware Architectures at Scale
by: Mayr, Matthias, et al.
Published: (2025)
by: Mayr, Matthias, et al.
Published: (2025)
Flexible Multi-Dimensional FFTs for Plane Wave Density Functional Theory Codes
by: Popovici, Doru Thom, et al.
Published: (2024)
by: Popovici, Doru Thom, et al.
Published: (2024)
HYLU: Hybrid Parallel Sparse LU Factorization
by: Chen, Xiaoming
Published: (2025)
by: Chen, Xiaoming
Published: (2025)
High-Performance Tensor Contraction without Transposition
by: Matthews, Devin A.
Published: (2016)
by: Matthews, Devin A.
Published: (2016)
Direct Low-Dose CT Image Reconstruction on GPU using Out-Of-Core: Precision and Quality Study
by: Chillarón, M., et al.
Published: (2024)
by: Chillarón, M., et al.
Published: (2024)
Scalable Dual Coordinate Descent for Kernel Methods
by: Shao, Zishan, et al.
Published: (2024)
by: Shao, Zishan, et al.
Published: (2024)
Accurate complex Jacobi rotations
by: Novaković, Vedran
Published: (2023)
by: Novaković, Vedran
Published: (2023)
Stencil Computations on AMD and Nvidia Graphics Processors: Performance and Tuning Strategies
by: Pekkilä, Johannes, et al.
Published: (2024)
by: Pekkilä, Johannes, et al.
Published: (2024)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
A natural language framework for non-conforming hybrid polytopal methods in Gridap.jl
by: Manyer, Jordi, et al.
Published: (2026)
by: Manyer, Jordi, et al.
Published: (2026)
A Task Parallel Orthonormalization Multigrid Method For Multiphase Elliptic Problems
by: Toprak, Teoman, et al.
Published: (2025)
by: Toprak, Teoman, et al.
Published: (2025)
Sensor Placement for Tsunami Early Warning via Large-Scale Bayesian Optimal Experimental Design
by: Venkat, Sreeram, et al.
Published: (2026)
by: Venkat, Sreeram, et al.
Published: (2026)
Distributed-memory Algorithms for Sparse Matrix Permutation, Extraction, and Assignment
by: Hassani, Elaheh, et al.
Published: (2025)
by: Hassani, Elaheh, et al.
Published: (2025)
Communication-Efficient and Memory-Aware Parallel Bootstrapping using MPI
by: Zhang, Di
Published: (2025)
by: Zhang, Di
Published: (2025)
A Virtual Processor brings back the Free Lunch
by: Kutschbach, Haymo
Published: (2026)
by: Kutschbach, Haymo
Published: (2026)
A Hybrid Direct-Iterative Method for Solving KKT Linear Systems
by: Regev, Shaked, et al.
Published: (2021)
by: Regev, Shaked, et al.
Published: (2021)
Accelerated Spatio-Temporal Bayesian Modeling for Multivariate Gaussian Processes
by: Gaedke-Merzhäuser, Lisa, et al.
Published: (2025)
by: Gaedke-Merzhäuser, Lisa, et al.
Published: (2025)
Chebyshev Accelerated Subspace Eigensolver for Pseudo-hermitian Hamiltonians
by: Di Napoli, Edoardo, et al.
Published: (2026)
by: Di Napoli, Edoardo, et al.
Published: (2026)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017)
by: Cicirello, Vincent A.
Published: (2017)
MIRGE: An Array-Based Computational Framework for Scientific Computing
by: Diener, Matthias, et al.
Published: (2025)
by: Diener, Matthias, et al.
Published: (2025)
Arithmetical enhancements of the Kogbetliantz method for the SVD of order two
by: Novaković, Vedran
Published: (2024)
by: Novaković, Vedran
Published: (2024)
Scaling the memory wall using mixed-precision -- HPG-MxP on an exascale machine
by: Kashi, Aditya, et al.
Published: (2025)
by: Kashi, Aditya, et al.
Published: (2025)
Adaptive time step selection for Spectral Deferred Correction
by: Saupe, Thomas, et al.
Published: (2024)
by: Saupe, Thomas, et al.
Published: (2024)
Resilience Against Soft Faults through Adaptivity in Spectral Deferred Correction
by: Saupe, Thomas, et al.
Published: (2024)
by: Saupe, Thomas, et al.
Published: (2024)
On Advanced Monte Carlo Methods for Linear Algebra on Advanced Accelerator Architectures
by: Lebedev, Anton, et al.
Published: (2024)
by: Lebedev, Anton, et al.
Published: (2024)
nuGPR: GPU-Accelerated Gaussian Process Regression with Iterative Algorithms and Low-Rank Approximations
by: Zhao, Ziqi, et al.
Published: (2025)
by: Zhao, Ziqi, et al.
Published: (2025)
Nearest Neighbors GParareal: Improving Scalability of Gaussian Processes for Parallel-in-Time Solvers
by: Gattiglio, Guglielmo, et al.
Published: (2024)
by: Gattiglio, Guglielmo, et al.
Published: (2024)
A C++17 Thread Pool for High-Performance Scientific Computing
by: Shoshany, Barak
Published: (2021)
by: Shoshany, Barak
Published: (2021)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
Incremental Hierarchical Tucker Decomposition
by: Aksoy, Doruk, et al.
Published: (2024)
by: Aksoy, Doruk, et al.
Published: (2024)
Similar Items
-
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
by: Kulkarni, Kaushik, et al.
Published: (2025) -
Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
by: Martinez-Ferrer, Pedro J., et al.
Published: (2025) -
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025) -
Bandicoot: A Templated C++ Library for GPU Linear Algebra
by: Curtin, Ryan R., et al.
Published: (2025) -
Exceeding the Numerical and Performance Characteristics of IEEE-754 SGEMM with BFloat16 Tensor Cores on GPUs for Scientific Computing
by: Bayraktar, Harun, et al.
Published: (2026)