Exploiting nested task-parallelism in the $\mathcal{H}-LU$ factorization
Fuente:
arXiv
Salvato in:
| Autori principali: | Carratalá-Sáez, Rocío, Christophersen, Sven, Aliaga, José I., Beltran, Vicenç, Börm, Steffen, Quintana-Ortí, Enrique S. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2019
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Distributed Tomographic Reconstruction with Quantization
di: Miao, Runxuan, et al.
Pubblicazione: (2024)
di: Miao, Runxuan, et al.
Pubblicazione: (2024)
Adaptive multiplication of $\mathcal{H}^2$-matrices with block-relative error control
di: Börm, Steffen
Pubblicazione: (2024)
di: Börm, Steffen
Pubblicazione: (2024)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
di: Kriemann, Ronald
Pubblicazione: (2024)
di: Kriemann, Ronald
Pubblicazione: (2024)
$\mathcal{H}^2$-matrices for translation-invariant kernel functions
di: Börm, Steffen, et al.
Pubblicazione: (2022)
di: Börm, Steffen, et al.
Pubblicazione: (2022)
A new Dune grid for scalable dynamic adaptivity based on the p4est software library
di: Burstedde, Carsten, et al.
Pubblicazione: (2025)
di: Burstedde, Carsten, et al.
Pubblicazione: (2025)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
di: Ferreira, Inês, et al.
Pubblicazione: (2024)
di: Ferreira, Inês, et al.
Pubblicazione: (2024)
Adaptive multiplication of rank-structured matrices in linear complexity
di: Börm, Steffen
Pubblicazione: (2023)
di: Börm, Steffen
Pubblicazione: (2023)
A scalable high-order multigrid-FFT Poisson solver for unbounded domains on adaptive multiresolution grids
di: Poncelet, Gilles, et al.
Pubblicazione: (2025)
di: Poncelet, Gilles, et al.
Pubblicazione: (2025)
Efficient Parallel Scheduling for Sparse Triangular Solvers
di: Böhnlein, Toni, et al.
Pubblicazione: (2025)
di: Böhnlein, Toni, et al.
Pubblicazione: (2025)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
di: Lei, Jie, et al.
Pubblicazione: (2024)
di: Lei, Jie, et al.
Pubblicazione: (2024)
Parallel performance of shared memory parallel spectral deferred corrections
di: Freese, Philip, et al.
Pubblicazione: (2024)
di: Freese, Philip, et al.
Pubblicazione: (2024)
Synthesis of signal processing algorithms with constraints on minimal parallelism and memory space
di: Salishev, Sergey
Pubblicazione: (2025)
di: Salishev, Sergey
Pubblicazione: (2025)
Leveraging Teaching on Demand: Approaching HPC to Undergrads
di: Catalán, S., et al.
Pubblicazione: (2026)
di: Catalán, S., et al.
Pubblicazione: (2026)
A multigrid reduction framework for domains with symmetries
di: Alsalti-Baldellou, Àdel, et al.
Pubblicazione: (2024)
di: Alsalti-Baldellou, Àdel, et al.
Pubblicazione: (2024)
The Performance of Low-Synchronization Variants of Reorthogonalized Block Classical Gram--Schmidt
di: Carson, Erin, et al.
Pubblicazione: (2025)
di: Carson, Erin, et al.
Pubblicazione: (2025)
CLAIRE: Scalable GPU-Accelerated Algorithms for Diffeomorphic Image Registration in 3D
di: Mang, Andreas
Pubblicazione: (2024)
di: Mang, Andreas
Pubblicazione: (2024)
A task-based data-flow methodology for programming heterogeneous systems with multiple accelerator APIs
di: Boné, Aleix, et al.
Pubblicazione: (2026)
di: Boné, Aleix, et al.
Pubblicazione: (2026)
Parallel Gauss-Jordan Elimination and System Reduction for Efficient Circuit Simulation
di: Noveski, Filip, et al.
Pubblicazione: (2026)
di: Noveski, Filip, et al.
Pubblicazione: (2026)
Multistep schemes for solving backward stochastic differential equations on GPU
di: Kapllani, Lorenc, et al.
Pubblicazione: (2019)
di: Kapllani, Lorenc, et al.
Pubblicazione: (2019)
RandNet-Parareal: a time-parallel PDE solver using Random Neural Networks
di: Gattiglio, Guglielmo, et al.
Pubblicazione: (2024)
di: Gattiglio, Guglielmo, et al.
Pubblicazione: (2024)
DMRlib: Easy-coding and Efficient Resource Management for Job Malleability
di: Iserte, Sergio, et al.
Pubblicazione: (2026)
di: Iserte, Sergio, et al.
Pubblicazione: (2026)
GPU-Accelerated Algorithms for Process Mapping
di: Samoldekin, Petr, et al.
Pubblicazione: (2025)
di: Samoldekin, Petr, et al.
Pubblicazione: (2025)
GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
di: Yin, Jia, et al.
Pubblicazione: (2025)
di: Yin, Jia, et al.
Pubblicazione: (2025)
OPTIMUM-DERAM: Highly Consistent, Scalable, and Secure Multi-Object Memory using RLNC
di: Nicolaou, Nicolas, et al.
Pubblicazione: (2026)
di: Nicolaou, Nicolas, et al.
Pubblicazione: (2026)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
di: Zhao, Haisha, et al.
Pubblicazione: (2025)
di: Zhao, Haisha, et al.
Pubblicazione: (2025)
Context Adaptive Cooperation
di: Albouy, Timothé, et al.
Pubblicazione: (2023)
di: Albouy, Timothé, et al.
Pubblicazione: (2023)
ML-Based Optimum Number of CUDA Streams for the GPU Implementation of the Tridiagonal Partition Method
di: Veneva, Milena, et al.
Pubblicazione: (2025)
di: Veneva, Milena, et al.
Pubblicazione: (2025)
ML-Based Optimum Sub-system Size Heuristic for the GPU Implementation of the Tridiagonal Partition Method
di: Veneva, Milena
Pubblicazione: (2025)
di: Veneva, Milena
Pubblicazione: (2025)
On efficient block Krylov-solvers for $\mathcal H^2$-matrices
di: Christophersen, Sven
Pubblicazione: (2025)
di: Christophersen, Sven
Pubblicazione: (2025)
Population Protocols Revisited: Parity and Beyond
di: Gąsieniec, Leszek, et al.
Pubblicazione: (2025)
di: Gąsieniec, Leszek, et al.
Pubblicazione: (2025)
Rethinking Thread Scheduling under Oversubscription: A User-Space Framework for Coordinating Multi-runtime and Multi-process Workloads
di: Roca, Aleix, et al.
Pubblicazione: (2026)
di: Roca, Aleix, et al.
Pubblicazione: (2026)
Direct Low-Dose CT Image Reconstruction on GPU using Out-Of-Core: Precision and Quality Study
di: Chillarón, M., et al.
Pubblicazione: (2024)
di: Chillarón, M., et al.
Pubblicazione: (2024)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
di: Gao, Jianhua, et al.
Pubblicazione: (2024)
di: Gao, Jianhua, et al.
Pubblicazione: (2024)
A Hybrid Direct-Iterative Method for Solving KKT Linear Systems
di: Regev, Shaked, et al.
Pubblicazione: (2021)
di: Regev, Shaked, et al.
Pubblicazione: (2021)
Massively-Parallel Implementation of Inextensible Elastic Rods Using Inter-block GPU Synchronization
di: Korzeniowski, Przemyslaw, et al.
Pubblicazione: (2025)
di: Korzeniowski, Przemyslaw, et al.
Pubblicazione: (2025)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
di: Venkat, Sreeram, et al.
Pubblicazione: (2025)
di: Venkat, Sreeram, et al.
Pubblicazione: (2025)
Energy efficiency optimization of task-parallel codes on asymmetric architectures
di: Costero, Luis, et al.
Pubblicazione: (2024)
di: Costero, Luis, et al.
Pubblicazione: (2024)
Nearest Neighbors GParareal: Improving Scalability of Gaussian Processes for Parallel-in-Time Solvers
di: Gattiglio, Guglielmo, et al.
Pubblicazione: (2024)
di: Gattiglio, Guglielmo, et al.
Pubblicazione: (2024)
Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach
di: Xiaoteng, et al.
Pubblicazione: (2024)
di: Xiaoteng, et al.
Pubblicazione: (2024)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
di: Suzuki, Kengo, et al.
Pubblicazione: (2026)
di: Suzuki, Kengo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Distributed Tomographic Reconstruction with Quantization
di: Miao, Runxuan, et al.
Pubblicazione: (2024) -
Adaptive multiplication of $\mathcal{H}^2$-matrices with block-relative error control
di: Börm, Steffen
Pubblicazione: (2024) -
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
di: Kriemann, Ronald
Pubblicazione: (2024) -
$\mathcal{H}^2$-matrices for translation-invariant kernel functions
di: Börm, Steffen, et al.
Pubblicazione: (2022) -
A new Dune grid for scalable dynamic adaptivity based on the p4est software library
di: Burstedde, Carsten, et al.
Pubblicazione: (2025)