Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bellavita, Julian, Pasquali, Thomas, Martin, Laura Del Rio, Vella, Flavio, Guidi, Giulia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
cuVegas: Accelerate Multidimensional Monte Carlo Integration through a Parallelized CUDA-based Implementation of the VEGAS Enhanced Algorithm
von: Tolotti, Emiliano, et al.
Veröffentlicht: (2024)
von: Tolotti, Emiliano, et al.
Veröffentlicht: (2024)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025)
von: Li, Shiju, et al.
Veröffentlicht: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
von: Xie, Xi, et al.
Veröffentlicht: (2024)
von: Xie, Xi, et al.
Veröffentlicht: (2024)
Accelerating Maximal Biclique Enumeration on GPUs
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
An inherently parallel H2-ULV factorization for solving dense linear systems on GPUs
von: Ma, Qianxiang, et al.
Veröffentlicht: (2025)
von: Ma, Qianxiang, et al.
Veröffentlicht: (2025)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
High-Performance Sorting-Based k-mer Counting in Distributed Memory with Flexible Hybrid Parallelism
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
Accelerating high-order continuum kinetic plasma simulations using multiple GPUs
von: Ho, Andrew, et al.
Veröffentlicht: (2024)
von: Ho, Andrew, et al.
Veröffentlicht: (2024)
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
von: Graça, Miguel, et al.
Veröffentlicht: (2026)
von: Graça, Miguel, et al.
Veröffentlicht: (2026)
Cuckoo-GPU: Accelerating Cuckoo Filters on Modern GPUs
von: Dortmann, Tim, et al.
Veröffentlicht: (2026)
von: Dortmann, Tim, et al.
Veröffentlicht: (2026)
State of practice: evaluating GPU performance of state vector and tensor network methods
von: Vallero, Marzio, et al.
Veröffentlicht: (2024)
von: Vallero, Marzio, et al.
Veröffentlicht: (2024)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
von: Ringoot, Evelyne, et al.
Veröffentlicht: (2025)
Combining Performance and Productivity: Accelerating the Network Sensing Graph Challenge with GPUs and Commodity Data Science Software
von: Samsi, Siddharth, et al.
Veröffentlicht: (2025)
von: Samsi, Siddharth, et al.
Veröffentlicht: (2025)
High Performance Unstructured SpMM Computation Using Tensor Cores
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
von: Shah, Milan, et al.
Veröffentlicht: (2026)
von: Shah, Milan, et al.
Veröffentlicht: (2026)
torch-sla: Differentiable Sparse Linear Algebra with Adjoint Solvers and Sparse Tensor Parallelism for PyTorch
von: Chi, Mingyuan, et al.
Veröffentlicht: (2026)
von: Chi, Mingyuan, et al.
Veröffentlicht: (2026)
GPU Accelerated Sparse Cholesky Factorization
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
von: Abubaker, Nabil, et al.
Veröffentlicht: (2024)
von: Abubaker, Nabil, et al.
Veröffentlicht: (2024)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
Accelerating Sparse DNNs Based on Tiled GEMM
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
Optimizing sDTW for AMD GPUs
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
von: Chang, Li-Wen, et al.
Veröffentlicht: (2024)
von: Chang, Li-Wen, et al.
Veröffentlicht: (2024)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
von: Bellavita, Julian, et al.
Veröffentlicht: (2026) -
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
von: Bellavita, Julian, et al.
Veröffentlicht: (2026) -
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
von: McFarland, Thomas, et al.
Veröffentlicht: (2025) -
cuVegas: Accelerate Multidimensional Monte Carlo Integration through a Parallelized CUDA-based Implementation of the VEGAS Enhanced Algorithm
von: Tolotti, Emiliano, et al.
Veröffentlicht: (2024) -
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)