GPU Implementations for Midsize Integer Addition and Multiplication
Fuente:
arXiv
Guardado en:
| Autores principales: | Oancea, Cosmin E., Watt, Stephen M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
por: Carrica, Vicki, et al.
Publicado: (2025)
por: Carrica, Vicki, et al.
Publicado: (2025)
Compiler-supported reduced precision and AoS-SoA transformations for heterogeneous hardware
por: Radtke, Pawel K., et al.
Publicado: (2025)
por: Radtke, Pawel K., et al.
Publicado: (2025)
Julia GraphBLAS with Nonblocking Execution
por: Costanza, Pascal, et al.
Publicado: (2025)
por: Costanza, Pascal, et al.
Publicado: (2025)
Verifying Properties of Index Arrays in a Purely-Functional Data-Parallel Language
por: Hinnerskov, Nikolaj Hey, et al.
Publicado: (2025)
por: Hinnerskov, Nikolaj Hey, et al.
Publicado: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
por: Villalobos, Johansell, et al.
Publicado: (2025)
por: Villalobos, Johansell, et al.
Publicado: (2025)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
por: Li, Yifan, et al.
Publicado: (2026)
por: Li, Yifan, et al.
Publicado: (2026)
LEGO: A Layout Expression Language for Code Generation of Hierarchical Mapping
por: Tavakkoli, Amir Mohammad, et al.
Publicado: (2025)
por: Tavakkoli, Amir Mohammad, et al.
Publicado: (2025)
Pipelined Dense Symmetric Eigenvalue Decomposition on Multi-GPU Architectures
por: Wang, Hansheng, et al.
Publicado: (2025)
por: Wang, Hansheng, et al.
Publicado: (2025)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
por: Bellavita, Julian, et al.
Publicado: (2026)
por: Bellavita, Julian, et al.
Publicado: (2026)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
por: Singh, Abhinav, et al.
Publicado: (2023)
por: Singh, Abhinav, et al.
Publicado: (2023)
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
por: Ringoot, Evelyne, et al.
Publicado: (2025)
por: Ringoot, Evelyne, et al.
Publicado: (2025)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
por: Zhu, Honglin, et al.
Publicado: (2026)
por: Zhu, Honglin, et al.
Publicado: (2026)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
por: Bernaschi, Massimo, et al.
Publicado: (2025)
por: Bernaschi, Massimo, et al.
Publicado: (2025)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
por: Panova, Elena, et al.
Publicado: (2022)
por: Panova, Elena, et al.
Publicado: (2022)
Scheduling Languages: A Past, Present, and Future Taxonomy
por: Hall, Mary, et al.
Publicado: (2024)
por: Hall, Mary, et al.
Publicado: (2024)
SPUMA: a minimally invasive approach to the GPU porting of OPENFOAM
por: Bnà, Simone, et al.
Publicado: (2025)
por: Bnà, Simone, et al.
Publicado: (2025)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
por: Llorente-Saguer, Isaac
Publicado: (2026)
por: Llorente-Saguer, Isaac
Publicado: (2026)
A shared compilation stack for distributed-memory parallelism in stencil DSLs
por: Bisbas, George, et al.
Publicado: (2024)
por: Bisbas, George, et al.
Publicado: (2024)
Robustness and Accuracy in Pipelined Bi-Conjugate Gradient Stabilized Method: A Comparative Study
por: Havdiak, Mykhailo, et al.
Publicado: (2024)
por: Havdiak, Mykhailo, et al.
Publicado: (2024)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
por: Ham, David A., et al.
Publicado: (2024)
por: Ham, David A., et al.
Publicado: (2024)
Enabling MPI communication within Numba/LLVM JIT-compiled Python code using numba-mpi v1.0
por: Derlatka, Kacper, et al.
Publicado: (2024)
por: Derlatka, Kacper, et al.
Publicado: (2024)
Enabling mixed-precision in spectral element codes
por: Chen, Yanxiang, et al.
Publicado: (2025)
por: Chen, Yanxiang, et al.
Publicado: (2025)
High-Performance Star-M SVD for Big Data Compression
por: Hussain, Md Taufique, et al.
Publicado: (2026)
por: Hussain, Md Taufique, et al.
Publicado: (2026)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
por: Ringoot, Evelyne, et al.
Publicado: (2025)
por: Ringoot, Evelyne, et al.
Publicado: (2025)
On the Challenges of Energy-Efficiency Analysis in HPC Systems: Evaluating Synthetic Benchmarks and Gromacs
por: Machado, Rafael Ravedutti Lucio, et al.
Publicado: (2025)
por: Machado, Rafael Ravedutti Lucio, et al.
Publicado: (2025)
A new open source framework for multiscale modeling of fibrous materials on heterogeneous supercomputers
por: Merson, Jacob, et al.
Publicado: (2023)
por: Merson, Jacob, et al.
Publicado: (2023)
SYCL compute kernels for ExaHyPE
por: Loi, Chung Ming, et al.
Publicado: (2023)
por: Loi, Chung Ming, et al.
Publicado: (2023)
Fray: An Efficient General-Purpose Concurrency Testing Platform for the JVM (Extended Version)
por: Li, Ao, et al.
Publicado: (2025)
por: Li, Ao, et al.
Publicado: (2025)
Determinacy with Priorities up to Clocks
por: Liquori, Luigi, et al.
Publicado: (2026)
por: Liquori, Luigi, et al.
Publicado: (2026)
GPU Accelerated Newton for Taylor Series Solutions of Polynomial Homotopies in Multiple Double Precision
por: Verschelde, Jan
Publicado: (2023)
por: Verschelde, Jan
Publicado: (2023)
PETSc/TAO Developments for GPU-Based Early Exascale Systems
por: Mills, Richard Tran, et al.
Publicado: (2024)
por: Mills, Richard Tran, et al.
Publicado: (2024)
Enabling mixed-precision with the help of tools: A Nekbone case study
por: Chen, Yanxiang, et al.
Publicado: (2024)
por: Chen, Yanxiang, et al.
Publicado: (2024)
Comparative analysis of large data processing in Apache Spark using Java, Python and Scala
por: Borodii, Ivan, et al.
Publicado: (2025)
por: Borodii, Ivan, et al.
Publicado: (2025)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
A Communication Avoiding and Reducing Algorithm for Symmetric Eigenproblem for Very Small Matrices
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
Automated MPI-X code generation for scalable finite-difference solvers
por: Bisbas, George, et al.
Publicado: (2023)
por: Bisbas, George, et al.
Publicado: (2023)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
por: Woller, Fabian, et al.
Publicado: (2025)
por: Woller, Fabian, et al.
Publicado: (2025)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
Performance measurements of modern Fortran MPI applications with Score-P
por: Corbin, Gregor
Publicado: (2025)
por: Corbin, Gregor
Publicado: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
por: Tu, Jiqun, et al.
Publicado: (2026)
por: Tu, Jiqun, et al.
Publicado: (2026)
Ejemplares similares
-
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
por: Carrica, Vicki, et al.
Publicado: (2025) -
Compiler-supported reduced precision and AoS-SoA transformations for heterogeneous hardware
por: Radtke, Pawel K., et al.
Publicado: (2025) -
Julia GraphBLAS with Nonblocking Execution
por: Costanza, Pascal, et al.
Publicado: (2025) -
Verifying Properties of Index Arrays in a Purely-Functional Data-Parallel Language
por: Hinnerskov, Nikolaj Hey, et al.
Publicado: (2025) -
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
por: Villalobos, Johansell, et al.
Publicado: (2025)