Matrix-Free Finite Volume Kernels on a Dataflow Architecture
Fuente:
arXiv
Salvato in:
| Autori principali: | Sai, Ryuichi, Hamon, Francois P., Mellor-Crummey, John, Araya-Polo, Mauricio |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
di: Sai, Ryuichi, et al.
Pubblicazione: (2023)
di: Sai, Ryuichi, et al.
Pubblicazione: (2023)
Giga-scale Kernel Matrix Vector Multiplication on GPU
di: Hu, Robert, et al.
Pubblicazione: (2022)
di: Hu, Robert, et al.
Pubblicazione: (2022)
Towards a Higher Roofline for Matrix-Vector Multiplication in Matrix-Free HOSFEM
di: Cao, Zijian, et al.
Pubblicazione: (2025)
di: Cao, Zijian, et al.
Pubblicazione: (2025)
The Software Landscape for the Density Matrix Renormalization Group
di: Sehlstedt, Per, et al.
Pubblicazione: (2025)
di: Sehlstedt, Per, et al.
Pubblicazione: (2025)
Rapid Variable Resolution Particle Initialization for Complex Geometries
di: Villodi, Navaneet, et al.
Pubblicazione: (2025)
di: Villodi, Navaneet, et al.
Pubblicazione: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
di: Xia, Yuning, et al.
Pubblicazione: (2026)
di: Xia, Yuning, et al.
Pubblicazione: (2026)
Methods for Few-View CT Image Reconstruction
di: Champley, Kyle M., et al.
Pubblicazione: (2024)
di: Champley, Kyle M., et al.
Pubblicazione: (2024)
trainsum -- A Python package for quantics tensor trains
di: Haubenwallner, Paul, et al.
Pubblicazione: (2026)
di: Haubenwallner, Paul, et al.
Pubblicazione: (2026)
Implementation of McMurchie-Davidson algorithm for Gaussian AO integrals suited for SIMD processors
di: Asadchev, Andrey, et al.
Pubblicazione: (2025)
di: Asadchev, Andrey, et al.
Pubblicazione: (2025)
Memory-Efficient Recursive Evaluation of 3-Center Gaussian Integrals
di: Asadchev, Andrey, et al.
Pubblicazione: (2022)
di: Asadchev, Andrey, et al.
Pubblicazione: (2022)
Recent Extensions of the ZKCM Library for Parallel and Accurate MPS Simulation of Quantum Circuits
di: SaiToh, Akira
Pubblicazione: (2024)
di: SaiToh, Akira
Pubblicazione: (2024)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
di: Li, Yifan, et al.
Pubblicazione: (2026)
di: Li, Yifan, et al.
Pubblicazione: (2026)
LeanBET: Formally-verified surface area calculations in Lean
di: Ugwuanyi, Ejike D., et al.
Pubblicazione: (2026)
di: Ugwuanyi, Ejike D., et al.
Pubblicazione: (2026)
A Performance Portable Matrix Free Dense MTTKRP in GenTen
di: Kosmacher, Gabriel, et al.
Pubblicazione: (2025)
di: Kosmacher, Gabriel, et al.
Pubblicazione: (2025)
Welding R and C++: A Tale of Two Programming Languages
di: Sepulveda, Mauricio Vargas
Pubblicazione: (2024)
di: Sepulveda, Mauricio Vargas
Pubblicazione: (2024)
f4ncgb: High Performance Gröbner Basis Computations in Free Algebras
di: Heisinger, Maximilian, et al.
Pubblicazione: (2025)
di: Heisinger, Maximilian, et al.
Pubblicazione: (2025)
OpenACC offloading of the MFC compressible multiphase flow solver on AMD and NVIDIA GPUs
di: Wilfong, Benjamin, et al.
Pubblicazione: (2024)
di: Wilfong, Benjamin, et al.
Pubblicazione: (2024)
GenML: A Python Library to Generate the Mittag-Leffler Correlated Noise
di: Qu, Xiang, et al.
Pubblicazione: (2024)
di: Qu, Xiang, et al.
Pubblicazione: (2024)
Hyper-reduction methods for accelerating nonlinear finite element simulations: open source implementation and reproducible benchmarks
di: Larsson, Axel, et al.
Pubblicazione: (2026)
di: Larsson, Axel, et al.
Pubblicazione: (2026)
Multi-GPU fast Fourier transforms in MATLAB (for large-scale phase-field crystal simulations)
di: Punke, Maik, et al.
Pubblicazione: (2026)
di: Punke, Maik, et al.
Pubblicazione: (2026)
Deriving Algorithms for Triangular Tridiagonalization a Skew-Symmetric Matrix
di: van de Geijn, Robert, et al.
Pubblicazione: (2023)
di: van de Geijn, Robert, et al.
Pubblicazione: (2023)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
di: Ringoot, Evelyne, et al.
Pubblicazione: (2025)
di: Ringoot, Evelyne, et al.
Pubblicazione: (2025)
Odd but Error-Free FastTwoSum: More General Conditions for FastTwoSum as an Error-Free Transformation for Faithful Rounding Modes
di: Park, Sehyeok, et al.
Pubblicazione: (2026)
di: Park, Sehyeok, et al.
Pubblicazione: (2026)
KHRONOS: a Kernel-Based Neural Architecture for Rapid, Resource-Efficient Scientific Computation
di: Batley, Reza T., et al.
Pubblicazione: (2025)
di: Batley, Reza T., et al.
Pubblicazione: (2025)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
di: Ham, David A., et al.
Pubblicazione: (2024)
di: Ham, David A., et al.
Pubblicazione: (2024)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
di: Zhu, Honglin, et al.
Pubblicazione: (2026)
di: Zhu, Honglin, et al.
Pubblicazione: (2026)
Harnessing Batched BLAS/LAPACK Kernels on GPUs for Parallel Solutions of Block Tridiagonal Systems
di: Jin, David, et al.
Pubblicazione: (2025)
di: Jin, David, et al.
Pubblicazione: (2025)
A Practical GPU-Enhanced Matrix-Free Primal-Dual Method for Large-Scale Conic Programs
di: Lin, Zhenwei, et al.
Pubblicazione: (2025)
di: Lin, Zhenwei, et al.
Pubblicazione: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
di: Tu, Jiqun, et al.
Pubblicazione: (2026)
di: Tu, Jiqun, et al.
Pubblicazione: (2026)
A Constraint-based Mathematical Modeling Library in Prolog with Answer Constraint Semantics
di: Fages, François
Pubblicazione: (2024)
di: Fages, François
Pubblicazione: (2024)
Sphractal: Estimating the Fractal Dimension of Surfaces Computed from Precise Atomic Coordinates via Box-Counting Algorithm
di: Ting, Jonathan Yik Chang, et al.
Pubblicazione: (2024)
di: Ting, Jonathan Yik Chang, et al.
Pubblicazione: (2024)
Hiperwalk: Simulation of Quantum Walks with Heterogeneous High-Performance Computing
di: Motta, Paulo, et al.
Pubblicazione: (2024)
di: Motta, Paulo, et al.
Pubblicazione: (2024)
SeQuant Framework for Symbolic and Numerical Tensor Algebra. I. Core Capabilities
di: Gaudel, Bimal, et al.
Pubblicazione: (2025)
di: Gaudel, Bimal, et al.
Pubblicazione: (2025)
Maestro: Intelligent Execution for Quantum Circuit Simulation
di: Bertomeu, Oriol, et al.
Pubblicazione: (2025)
di: Bertomeu, Oriol, et al.
Pubblicazione: (2025)
Performance Analysis of Effective Symbolic Methods for Solving Band Matrix SLAEs
di: Veneva, Milena, et al.
Pubblicazione: (2019)
di: Veneva, Milena, et al.
Pubblicazione: (2019)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
di: Li, Junjie
Pubblicazione: (2024)
di: Li, Junjie
Pubblicazione: (2024)
GeoWarp: An automatically differentiable and GPU-accelerated implicit MPM framework for geomechanics based on NVIDIA Warp
di: Zhao, Yidong, et al.
Pubblicazione: (2025)
di: Zhao, Yidong, et al.
Pubblicazione: (2025)
Large-Scale Simulations of Turbulent Flows using Lattice Boltzmann Methods on Heterogeneous High Performance Computers
di: Kummerländer, Adrian, et al.
Pubblicazione: (2025)
di: Kummerländer, Adrian, et al.
Pubblicazione: (2025)
Unlocking massively parallel spectral proper orthogonal decompositions in the PySPOD package
di: Rogowski, Marcin, et al.
Pubblicazione: (2023)
di: Rogowski, Marcin, et al.
Pubblicazione: (2023)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
di: Wang, Hansheng, et al.
Pubblicazione: (2025)
di: Wang, Hansheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
di: Sai, Ryuichi, et al.
Pubblicazione: (2023) -
Giga-scale Kernel Matrix Vector Multiplication on GPU
di: Hu, Robert, et al.
Pubblicazione: (2022) -
Towards a Higher Roofline for Matrix-Vector Multiplication in Matrix-Free HOSFEM
di: Cao, Zijian, et al.
Pubblicazione: (2025) -
The Software Landscape for the Density Matrix Renormalization Group
di: Sehlstedt, Per, et al.
Pubblicazione: (2025) -
Rapid Variable Resolution Particle Initialization for Complex Geometries
di: Villodi, Navaneet, et al.
Pubblicazione: (2025)