Matrix-Free Finite Volume Kernels on a Dataflow Architecture
Fuente:
arXiv
Guardado en:
| Autores principales: | Sai, Ryuichi, Hamon, Francois P., Mellor-Crummey, John, Araya-Polo, Mauricio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
por: Sai, Ryuichi, et al.
Publicado: (2023)
por: Sai, Ryuichi, et al.
Publicado: (2023)
Giga-scale Kernel Matrix Vector Multiplication on GPU
por: Hu, Robert, et al.
Publicado: (2022)
por: Hu, Robert, et al.
Publicado: (2022)
Towards a Higher Roofline for Matrix-Vector Multiplication in Matrix-Free HOSFEM
por: Cao, Zijian, et al.
Publicado: (2025)
por: Cao, Zijian, et al.
Publicado: (2025)
The Software Landscape for the Density Matrix Renormalization Group
por: Sehlstedt, Per, et al.
Publicado: (2025)
por: Sehlstedt, Per, et al.
Publicado: (2025)
Rapid Variable Resolution Particle Initialization for Complex Geometries
por: Villodi, Navaneet, et al.
Publicado: (2025)
por: Villodi, Navaneet, et al.
Publicado: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
por: Xia, Yuning, et al.
Publicado: (2026)
por: Xia, Yuning, et al.
Publicado: (2026)
Methods for Few-View CT Image Reconstruction
por: Champley, Kyle M., et al.
Publicado: (2024)
por: Champley, Kyle M., et al.
Publicado: (2024)
trainsum -- A Python package for quantics tensor trains
por: Haubenwallner, Paul, et al.
Publicado: (2026)
por: Haubenwallner, Paul, et al.
Publicado: (2026)
Implementation of McMurchie-Davidson algorithm for Gaussian AO integrals suited for SIMD processors
por: Asadchev, Andrey, et al.
Publicado: (2025)
por: Asadchev, Andrey, et al.
Publicado: (2025)
Memory-Efficient Recursive Evaluation of 3-Center Gaussian Integrals
por: Asadchev, Andrey, et al.
Publicado: (2022)
por: Asadchev, Andrey, et al.
Publicado: (2022)
Recent Extensions of the ZKCM Library for Parallel and Accurate MPS Simulation of Quantum Circuits
por: SaiToh, Akira
Publicado: (2024)
por: SaiToh, Akira
Publicado: (2024)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
por: Li, Yifan, et al.
Publicado: (2026)
por: Li, Yifan, et al.
Publicado: (2026)
LeanBET: Formally-verified surface area calculations in Lean
por: Ugwuanyi, Ejike D., et al.
Publicado: (2026)
por: Ugwuanyi, Ejike D., et al.
Publicado: (2026)
A Performance Portable Matrix Free Dense MTTKRP in GenTen
por: Kosmacher, Gabriel, et al.
Publicado: (2025)
por: Kosmacher, Gabriel, et al.
Publicado: (2025)
Welding R and C++: A Tale of Two Programming Languages
por: Sepulveda, Mauricio Vargas
Publicado: (2024)
por: Sepulveda, Mauricio Vargas
Publicado: (2024)
f4ncgb: High Performance Gröbner Basis Computations in Free Algebras
por: Heisinger, Maximilian, et al.
Publicado: (2025)
por: Heisinger, Maximilian, et al.
Publicado: (2025)
OpenACC offloading of the MFC compressible multiphase flow solver on AMD and NVIDIA GPUs
por: Wilfong, Benjamin, et al.
Publicado: (2024)
por: Wilfong, Benjamin, et al.
Publicado: (2024)
GenML: A Python Library to Generate the Mittag-Leffler Correlated Noise
por: Qu, Xiang, et al.
Publicado: (2024)
por: Qu, Xiang, et al.
Publicado: (2024)
Hyper-reduction methods for accelerating nonlinear finite element simulations: open source implementation and reproducible benchmarks
por: Larsson, Axel, et al.
Publicado: (2026)
por: Larsson, Axel, et al.
Publicado: (2026)
Multi-GPU fast Fourier transforms in MATLAB (for large-scale phase-field crystal simulations)
por: Punke, Maik, et al.
Publicado: (2026)
por: Punke, Maik, et al.
Publicado: (2026)
Deriving Algorithms for Triangular Tridiagonalization a Skew-Symmetric Matrix
por: van de Geijn, Robert, et al.
Publicado: (2023)
por: van de Geijn, Robert, et al.
Publicado: (2023)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
por: Ringoot, Evelyne, et al.
Publicado: (2025)
por: Ringoot, Evelyne, et al.
Publicado: (2025)
Odd but Error-Free FastTwoSum: More General Conditions for FastTwoSum as an Error-Free Transformation for Faithful Rounding Modes
por: Park, Sehyeok, et al.
Publicado: (2026)
por: Park, Sehyeok, et al.
Publicado: (2026)
KHRONOS: a Kernel-Based Neural Architecture for Rapid, Resource-Efficient Scientific Computation
por: Batley, Reza T., et al.
Publicado: (2025)
por: Batley, Reza T., et al.
Publicado: (2025)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
por: Ham, David A., et al.
Publicado: (2024)
por: Ham, David A., et al.
Publicado: (2024)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
por: Zhu, Honglin, et al.
Publicado: (2026)
por: Zhu, Honglin, et al.
Publicado: (2026)
Harnessing Batched BLAS/LAPACK Kernels on GPUs for Parallel Solutions of Block Tridiagonal Systems
por: Jin, David, et al.
Publicado: (2025)
por: Jin, David, et al.
Publicado: (2025)
A Practical GPU-Enhanced Matrix-Free Primal-Dual Method for Large-Scale Conic Programs
por: Lin, Zhenwei, et al.
Publicado: (2025)
por: Lin, Zhenwei, et al.
Publicado: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
por: Tu, Jiqun, et al.
Publicado: (2026)
por: Tu, Jiqun, et al.
Publicado: (2026)
A Constraint-based Mathematical Modeling Library in Prolog with Answer Constraint Semantics
por: Fages, François
Publicado: (2024)
por: Fages, François
Publicado: (2024)
Sphractal: Estimating the Fractal Dimension of Surfaces Computed from Precise Atomic Coordinates via Box-Counting Algorithm
por: Ting, Jonathan Yik Chang, et al.
Publicado: (2024)
por: Ting, Jonathan Yik Chang, et al.
Publicado: (2024)
Hiperwalk: Simulation of Quantum Walks with Heterogeneous High-Performance Computing
por: Motta, Paulo, et al.
Publicado: (2024)
por: Motta, Paulo, et al.
Publicado: (2024)
SeQuant Framework for Symbolic and Numerical Tensor Algebra. I. Core Capabilities
por: Gaudel, Bimal, et al.
Publicado: (2025)
por: Gaudel, Bimal, et al.
Publicado: (2025)
Maestro: Intelligent Execution for Quantum Circuit Simulation
por: Bertomeu, Oriol, et al.
Publicado: (2025)
por: Bertomeu, Oriol, et al.
Publicado: (2025)
Performance Analysis of Effective Symbolic Methods for Solving Band Matrix SLAEs
por: Veneva, Milena, et al.
Publicado: (2019)
por: Veneva, Milena, et al.
Publicado: (2019)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
por: Li, Junjie
Publicado: (2024)
por: Li, Junjie
Publicado: (2024)
GeoWarp: An automatically differentiable and GPU-accelerated implicit MPM framework for geomechanics based on NVIDIA Warp
por: Zhao, Yidong, et al.
Publicado: (2025)
por: Zhao, Yidong, et al.
Publicado: (2025)
Large-Scale Simulations of Turbulent Flows using Lattice Boltzmann Methods on Heterogeneous High Performance Computers
por: Kummerländer, Adrian, et al.
Publicado: (2025)
por: Kummerländer, Adrian, et al.
Publicado: (2025)
Unlocking massively parallel spectral proper orthogonal decompositions in the PySPOD package
por: Rogowski, Marcin, et al.
Publicado: (2023)
por: Rogowski, Marcin, et al.
Publicado: (2023)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
por: Wang, Hansheng, et al.
Publicado: (2025)
por: Wang, Hansheng, et al.
Publicado: (2025)
Ejemplares similares
-
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
por: Sai, Ryuichi, et al.
Publicado: (2023) -
Giga-scale Kernel Matrix Vector Multiplication on GPU
por: Hu, Robert, et al.
Publicado: (2022) -
Towards a Higher Roofline for Matrix-Vector Multiplication in Matrix-Free HOSFEM
por: Cao, Zijian, et al.
Publicado: (2025) -
The Software Landscape for the Density Matrix Renormalization Group
por: Sehlstedt, Per, et al.
Publicado: (2025) -
Rapid Variable Resolution Particle Initialization for Complex Geometries
por: Villodi, Navaneet, et al.
Publicado: (2025)