ML-Based Optimum Number of CUDA Streams for the GPU Implementation of the Tridiagonal Partition Method
Fuente:
arXiv
Guardado en:
| Autores principales: | Veneva, Milena, Imamura, Toshiyuki |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ML-Based Optimum Sub-system Size Heuristic for the GPU Implementation of the Tridiagonal Partition Method
por: Veneva, Milena
Publicado: (2025)
por: Veneva, Milena
Publicado: (2025)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
por: Kriemann, Ronald
Publicado: (2024)
por: Kriemann, Ronald
Publicado: (2024)
Parallel Gauss-Jordan Elimination and System Reduction for Efficient Circuit Simulation
por: Noveski, Filip, et al.
Publicado: (2026)
por: Noveski, Filip, et al.
Publicado: (2026)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
por: Venkat, Sreeram, et al.
Publicado: (2025)
por: Venkat, Sreeram, et al.
Publicado: (2025)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
por: Ferreira, Inês, et al.
Publicado: (2024)
por: Ferreira, Inês, et al.
Publicado: (2024)
Adaptive time step selection for Spectral Deferred Correction
por: Saupe, Thomas, et al.
Publicado: (2024)
por: Saupe, Thomas, et al.
Publicado: (2024)
Resilience Against Soft Faults through Adaptivity in Spectral Deferred Correction
por: Saupe, Thomas, et al.
Publicado: (2024)
por: Saupe, Thomas, et al.
Publicado: (2024)
CLAIRE: Scalable GPU-Accelerated Algorithms for Diffeomorphic Image Registration in 3D
por: Mang, Andreas
Publicado: (2024)
por: Mang, Andreas
Publicado: (2024)
A Task Parallel Orthonormalization Multigrid Method For Multiphase Elliptic Problems
por: Toprak, Teoman, et al.
Publicado: (2025)
por: Toprak, Teoman, et al.
Publicado: (2025)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
The Performance of Low-Synchronization Variants of Reorthogonalized Block Classical Gram--Schmidt
por: Carson, Erin, et al.
Publicado: (2025)
por: Carson, Erin, et al.
Publicado: (2025)
nuGPR: GPU-Accelerated Gaussian Process Regression with Iterative Algorithms and Low-Rank Approximations
por: Zhao, Ziqi, et al.
Publicado: (2025)
por: Zhao, Ziqi, et al.
Publicado: (2025)
High-performance matrix-free unfitted finite element operator evaluation
por: Bergbauer, Maximilian, et al.
Publicado: (2024)
por: Bergbauer, Maximilian, et al.
Publicado: (2024)
Matrix-Free Evaluation of High-Order Shifted Boundary Finite Element Operators
por: Wichrowski, Michał
Publicado: (2025)
por: Wichrowski, Michał
Publicado: (2025)
Scalable Mean-Variance Portfolio Optimization via Subspace Embeddings and GPU-Friendly Nesterov-Accelerated Projected Gradient
por: Niu, Yi-Shuai, et al.
Publicado: (2026)
por: Niu, Yi-Shuai, et al.
Publicado: (2026)
RandNet-Parareal: a time-parallel PDE solver using Random Neural Networks
por: Gattiglio, Guglielmo, et al.
Publicado: (2024)
por: Gattiglio, Guglielmo, et al.
Publicado: (2024)
Symbolic Algorithm for Solving SLAEs with Multi-Diagonal Coefficient Matrices
por: Veneva, Milena
Publicado: (2024)
por: Veneva, Milena
Publicado: (2024)
M2L Translation Operators for Kernel Independent Fast Multipole Methods on Modern Architectures
por: Kailasa, Srinath, et al.
Publicado: (2024)
por: Kailasa, Srinath, et al.
Publicado: (2024)
A multigrid reduction framework for domains with symmetries
por: Alsalti-Baldellou, Àdel, et al.
Publicado: (2024)
por: Alsalti-Baldellou, Àdel, et al.
Publicado: (2024)
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
por: Kulkarni, Kaushik, et al.
Publicado: (2025)
por: Kulkarni, Kaushik, et al.
Publicado: (2025)
Iterative Methods in GPU-Resident Linear Solvers for Nonlinear Constrained Optimization
por: Świrydowicz, Kasia, et al.
Publicado: (2024)
por: Świrydowicz, Kasia, et al.
Publicado: (2024)
A Proximal-Gradient Method for Constrained Optimization
por: Dai, Yutong, et al.
Publicado: (2024)
por: Dai, Yutong, et al.
Publicado: (2024)
A multigrid method for CutFEM and its implementation on GPU
por: Cui, Cu, et al.
Publicado: (2025)
por: Cui, Cu, et al.
Publicado: (2025)
Accelerated primal dual fixed point algorithm
por: Zhu, Ya-Nan
Publicado: (2025)
por: Zhu, Ya-Nan
Publicado: (2025)
A Proximal-Gradient Method for Solving Regularized Optimization Problems with General Constraints
por: Curtis, Frank E., et al.
Publicado: (2025)
por: Curtis, Frank E., et al.
Publicado: (2025)
Nearest Neighbors GParareal: Improving Scalability of Gaussian Processes for Parallel-in-Time Solvers
por: Gattiglio, Guglielmo, et al.
Publicado: (2024)
por: Gattiglio, Guglielmo, et al.
Publicado: (2024)
Parallel performance of shared memory parallel spectral deferred corrections
por: Freese, Philip, et al.
Publicado: (2024)
por: Freese, Philip, et al.
Publicado: (2024)
Subspace-constrained randomized coordinate descent for linear systems with good low-rank matrix approximations
por: Lok, Jackie, et al.
Publicado: (2025)
por: Lok, Jackie, et al.
Publicado: (2025)
Scaling the memory wall using mixed-precision -- HPG-MxP on an exascale machine
por: Kashi, Aditya, et al.
Publicado: (2025)
por: Kashi, Aditya, et al.
Publicado: (2025)
Parallel-in-time Multilevel Krylov Methods: A Prototype
por: Erlangga, Yogi A.
Publicado: (2023)
por: Erlangga, Yogi A.
Publicado: (2023)
Prob-GParareal: A Probabilistic Numerical Parallel-in-Time Solver for Differential Equations
por: Gattiglio, Guglielmo, et al.
Publicado: (2025)
por: Gattiglio, Guglielmo, et al.
Publicado: (2025)
Scalable Multilevel Monte Carlo Methods Exploiting Parallel Redistribution on Coarse Levels
por: Fairbanks, Hillary R., et al.
Publicado: (2024)
por: Fairbanks, Hillary R., et al.
Publicado: (2024)
Small errors in random zeroth-order optimization are imaginary
por: Jongeneel, Wouter, et al.
Publicado: (2021)
por: Jongeneel, Wouter, et al.
Publicado: (2021)
Parametrization and convergence of a primal-dual block-coordinate approach to linearly-constrained nonsmooth optimization
por: Bilenne, Olivier
Publicado: (2024)
por: Bilenne, Olivier
Publicado: (2024)
On the Relationships among GPU-Accelerated First-Order Methods for Solving Linear Programming
por: Chen, Kaihuang, et al.
Publicado: (2025)
por: Chen, Kaihuang, et al.
Publicado: (2025)
Low-Memory Numerical Certification
por: Breiding, Paul, et al.
Publicado: (2026)
por: Breiding, Paul, et al.
Publicado: (2026)
Scalable Dual Coordinate Descent for Kernel Methods
por: Shao, Zishan, et al.
Publicado: (2024)
por: Shao, Zishan, et al.
Publicado: (2024)
A Parareal Algorithm with Low-Rank Coarse Solvers
por: Gander, Martin J., et al.
Publicado: (2025)
por: Gander, Martin J., et al.
Publicado: (2025)
Optimization in Theory and Practice
por: Wright, Stephen J.
Publicado: (2025)
por: Wright, Stephen J.
Publicado: (2025)
Ejemplares similares
-
ML-Based Optimum Sub-system Size Heuristic for the GPU Implementation of the Tridiagonal Partition Method
por: Veneva, Milena
Publicado: (2025) -
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
por: Kriemann, Ronald
Publicado: (2024) -
Parallel Gauss-Jordan Elimination and System Reduction for Efficient Circuit Simulation
por: Noveski, Filip, et al.
Publicado: (2026) -
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
por: Venkat, Sreeram, et al.
Publicado: (2025) -
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)