A High Performance GPU CountSketch Implementation and Its Application to Multisketching and Least Squares Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Higgins, Andrew J., Boman, Erik G., Yamazaki, Ichitaro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two-Stage Block Orthogonalization to Improve Performance of $s$-step GMRES
by: Yamazaki, Ichitaro, et al.
Published: (2024)
by: Yamazaki, Ichitaro, et al.
Published: (2024)
Random-sketching Techniques to Enhance the Numerical Stability of Block Orthogonalization Algorithms for s-step GMRES
by: Yamazaki, Ichitaro, et al.
Published: (2025)
by: Yamazaki, Ichitaro, et al.
Published: (2025)
Efficient and scalable atmospheric dynamics simulations using non-conforming meshes
by: Orlando, Giuseppe, et al.
Published: (2024)
by: Orlando, Giuseppe, et al.
Published: (2024)
Improving the scalability of a high-order atmospheric dynamics solver based on the deal.II library
by: Orlando, Giuseppe, et al.
Published: (2025)
by: Orlando, Giuseppe, et al.
Published: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
by: Pilliat, Emmanuel
Published: (2026)
by: Pilliat, Emmanuel
Published: (2026)
GPU-Parallelizable Randomized Sketch-and-Precondition for Linear Regression using Sparse Sign Sketches
by: Chen, Tyler, et al.
Published: (2025)
by: Chen, Tyler, et al.
Published: (2025)
Scalable GPU Performance Variability Analysis framework
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
Taking GPU Programming Models to Task for Performance Portability
by: Davis, Joshua H., et al.
Published: (2024)
by: Davis, Joshua H., et al.
Published: (2024)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
by: Davis, Joshua H., et al.
Published: (2026)
by: Davis, Joshua H., et al.
Published: (2026)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
Cucheb: A GPU implementation of the filtered Lanczos procedure
by: Aurentz, Jared L., et al.
Published: (2024)
by: Aurentz, Jared L., et al.
Published: (2024)
Towards a GPU-Parallelization of the neXtSIM-DG Dynamical Core
by: Jendersie, Robert, et al.
Published: (2024)
by: Jendersie, Robert, et al.
Published: (2024)
GPU Accelerated Implicit Kinetic Meshfree Method based on Modified LU-SGS
by: Verma, Mayuri, et al.
Published: (2024)
by: Verma, Mayuri, et al.
Published: (2024)
Serinv: A Scalable Library for the Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
by: Maillou, Vincent, et al.
Published: (2025)
by: Maillou, Vincent, et al.
Published: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
MPI Implementation Profiling for Better Application Performance
by: Shipley, Riley, et al.
Published: (2024)
by: Shipley, Riley, et al.
Published: (2024)
The Energy Cost of Execution-Idle in GPU Clusters
by: Lei, Yiran, et al.
Published: (2026)
by: Lei, Yiran, et al.
Published: (2026)
On the Partitioning of GPU Power among Multi-Instances
by: Vamja, Tirth, et al.
Published: (2025)
by: Vamja, Tirth, et al.
Published: (2025)
A simple GPU implementation of spectral-element methods for solving 3D Poisson type equations on rectangular domains and its applications
by: Liu, Xinyu, et al.
Published: (2023)
by: Liu, Xinyu, et al.
Published: (2023)
Disaggregated Design for GPU-Based Volumetric Data Structures
by: Meneghin, Massimiliano, et al.
Published: (2025)
by: Meneghin, Massimiliano, et al.
Published: (2025)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
by: Apanasevich, L., et al.
Published: (2024)
by: Apanasevich, L., et al.
Published: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
by: Zhao, Yanbo, et al.
Published: (2025)
by: Zhao, Yanbo, et al.
Published: (2025)
Denoising Application Performance Models with Noise-Resilient Priors
by: de Morais, Gustavo, et al.
Published: (2025)
by: de Morais, Gustavo, et al.
Published: (2025)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
by: Islam, Tanzima Z., et al.
Published: (2024)
by: Islam, Tanzima Z., et al.
Published: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
by: Liu, Shifang, et al.
Published: (2025)
by: Liu, Shifang, et al.
Published: (2025)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
by: Wang, Yidi, et al.
Published: (2024)
by: Wang, Yidi, et al.
Published: (2024)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
by: Chang, Dali, et al.
Published: (2026)
by: Chang, Dali, et al.
Published: (2026)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
by: Wahlgren, Jacob, et al.
Published: (2025)
by: Wahlgren, Jacob, et al.
Published: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
by: Xia, Yuning, et al.
Published: (2026)
by: Xia, Yuning, et al.
Published: (2026)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
by: Curless, Brian, et al.
Published: (2025)
by: Curless, Brian, et al.
Published: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
by: Lin, Mao, et al.
Published: (2026)
by: Lin, Mao, et al.
Published: (2026)
A Parallel in Time Algorithm Based on ParaExp for Optimal Control Problems
by: Kwok, Felix, et al.
Published: (2024)
by: Kwok, Felix, et al.
Published: (2024)
RAPTOR: Practical Numerical Profiling of Scientific Applications
by: Hoerold, Faveo, et al.
Published: (2025)
by: Hoerold, Faveo, et al.
Published: (2025)
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
by: Kulkarni, Kaushik, et al.
Published: (2025)
by: Kulkarni, Kaushik, et al.
Published: (2025)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
by: Kulkarni, Apurv Deepak, et al.
Published: (2025)
by: Kulkarni, Apurv Deepak, et al.
Published: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
by: Ramesh, Risshab Srinivas
Published: (2024)
by: Ramesh, Risshab Srinivas
Published: (2024)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
by: Accordi, Gianmarco, et al.
Published: (2025)
by: Accordi, Gianmarco, et al.
Published: (2025)
Similar Items
-
Two-Stage Block Orthogonalization to Improve Performance of $s$-step GMRES
by: Yamazaki, Ichitaro, et al.
Published: (2024) -
Random-sketching Techniques to Enhance the Numerical Stability of Block Orthogonalization Algorithms for s-step GMRES
by: Yamazaki, Ichitaro, et al.
Published: (2025) -
Efficient and scalable atmospheric dynamics simulations using non-conforming meshes
by: Orlando, Giuseppe, et al.
Published: (2024) -
Improving the scalability of a high-order atmospheric dynamics solver based on the deal.II library
by: Orlando, Giuseppe, et al.
Published: (2025) -
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
by: Chen, Qian, et al.
Published: (2024)