Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
Fuente:
arXiv
Guardado en:
| Autores principales: | Bellavita, Julian, Rubino, Matthew, Iyer, Nakul, Chang, Andrew, Devarakonda, Aditya, Vella, Flavio, Guidi, Giulia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
por: Bellavita, Julian, et al.
Publicado: (2025)
por: Bellavita, Julian, et al.
Publicado: (2025)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
por: Bellavita, Julian, et al.
Publicado: (2026)
por: Bellavita, Julian, et al.
Publicado: (2026)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
por: McFarland, Thomas, et al.
Publicado: (2025)
por: McFarland, Thomas, et al.
Publicado: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
por: Bhosale, Aditya, et al.
Publicado: (2025)
por: Bhosale, Aditya, et al.
Publicado: (2025)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
por: Chang, Li-Wen, et al.
Publicado: (2024)
por: Chang, Li-Wen, et al.
Publicado: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
por: Ekelund, Jonah, et al.
Publicado: (2025)
por: Ekelund, Jonah, et al.
Publicado: (2025)
cuVegas: Accelerate Multidimensional Monte Carlo Integration through a Parallelized CUDA-based Implementation of the VEGAS Enhanced Algorithm
por: Tolotti, Emiliano, et al.
Publicado: (2024)
por: Tolotti, Emiliano, et al.
Publicado: (2024)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
por: Dwaraknath, Rajat Vadiraj, et al.
Publicado: (2026)
por: Dwaraknath, Rajat Vadiraj, et al.
Publicado: (2026)
High-Performance Sorting-Based k-mer Counting in Distributed Memory with Flexible Hybrid Parallelism
por: Li, Yifan, et al.
Publicado: (2024)
por: Li, Yifan, et al.
Publicado: (2024)
Accelerating Maximal Biclique Enumeration on GPUs
por: Hsieh, Chou-Ying, et al.
Publicado: (2024)
por: Hsieh, Chou-Ying, et al.
Publicado: (2024)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
por: Li, Aiying, et al.
Publicado: (2026)
por: Li, Aiying, et al.
Publicado: (2026)
Parallelizing Maximal Clique Enumeration on GPUs
por: Almasri, Mohammad, et al.
Publicado: (2022)
por: Almasri, Mohammad, et al.
Publicado: (2022)
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
por: Xie, Xi, et al.
Publicado: (2024)
por: Xie, Xi, et al.
Publicado: (2024)
State of practice: evaluating GPU performance of state vector and tensor network methods
por: Vallero, Marzio, et al.
Publicado: (2024)
por: Vallero, Marzio, et al.
Publicado: (2024)
DistShap: Scalable GNN Explanations with Distributed Shapley Values
por: Akkas, Selahattin, et al.
Publicado: (2025)
por: Akkas, Selahattin, et al.
Publicado: (2025)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
por: Cui, Shengkun, et al.
Publicado: (2025)
por: Cui, Shengkun, et al.
Publicado: (2025)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
por: Li, Yifan, et al.
Publicado: (2026)
por: Li, Yifan, et al.
Publicado: (2026)
Accelerating high-order continuum kinetic plasma simulations using multiple GPUs
por: Ho, Andrew, et al.
Publicado: (2024)
por: Ho, Andrew, et al.
Publicado: (2024)
High Performance Unstructured SpMM Computation Using Tensor Cores
por: Okanovic, Patrik, et al.
Publicado: (2024)
por: Okanovic, Patrik, et al.
Publicado: (2024)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
por: Hu, Tiancheng, et al.
Publicado: (2026)
por: Hu, Tiancheng, et al.
Publicado: (2026)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
por: He, Guoliang, et al.
Publicado: (2025)
por: He, Guoliang, et al.
Publicado: (2025)
Optimizing sDTW for AMD GPUs
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
por: Tramm, John, et al.
Publicado: (2024)
por: Tramm, John, et al.
Publicado: (2024)
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
por: Pan, Qilong, et al.
Publicado: (2025)
por: Pan, Qilong, et al.
Publicado: (2025)
Scalable Dual Coordinate Descent for Kernel Methods
por: Shao, Zishan, et al.
Publicado: (2024)
por: Shao, Zishan, et al.
Publicado: (2024)
Serving Compound Inference Systems on Datacenter GPUs
por: Devata, Sriram, et al.
Publicado: (2026)
por: Devata, Sriram, et al.
Publicado: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
por: Jangda, Abhinav, et al.
Publicado: (2024)
por: Jangda, Abhinav, et al.
Publicado: (2024)
Optimal Workload Placement on Multi-Instance GPUs
por: Turkkan, Bekir, et al.
Publicado: (2024)
por: Turkkan, Bekir, et al.
Publicado: (2024)
High-Performance Star-M SVD for Big Data Compression
por: Hussain, Md Taufique, et al.
Publicado: (2026)
por: Hussain, Md Taufique, et al.
Publicado: (2026)
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
por: Zheng, Size, et al.
Publicado: (2025)
por: Zheng, Size, et al.
Publicado: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
por: Zhang, Zeyu, et al.
Publicado: (2025)
por: Zhang, Zeyu, et al.
Publicado: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
por: Brock, Benjamin, et al.
Publicado: (2023)
por: Brock, Benjamin, et al.
Publicado: (2023)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
por: Plesner, Andreas, et al.
Publicado: (2024)
por: Plesner, Andreas, et al.
Publicado: (2024)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
por: Yang, Shuo, et al.
Publicado: (2026)
por: Yang, Shuo, et al.
Publicado: (2026)
Managing Multi Instance GPUs for High Throughput and Energy Savings
por: Saraha, Abhijeet, et al.
Publicado: (2025)
por: Saraha, Abhijeet, et al.
Publicado: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
por: Ernst, Dominik, et al.
Publicado: (2022)
por: Ernst, Dominik, et al.
Publicado: (2022)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
Quantum-Enhanced Distributed Sensor Fusion: Lower Bounds on Aggregation from Projection Noise to Heisenberg-Limited Byzantine-Tolerant Networks
por: Iyer, Vasanth, et al.
Publicado: (2026)
por: Iyer, Vasanth, et al.
Publicado: (2026)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
por: Gao, Wei, et al.
Publicado: (2026)
por: Gao, Wei, et al.
Publicado: (2026)
Anonymized Network Sensing using C++26 std::execution on GPUs
por: Mandulak, Michael, et al.
Publicado: (2025)
por: Mandulak, Michael, et al.
Publicado: (2025)
Ejemplares similares
-
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
por: Bellavita, Julian, et al.
Publicado: (2025) -
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
por: Bellavita, Julian, et al.
Publicado: (2026) -
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
por: McFarland, Thomas, et al.
Publicado: (2025) -
An Adaptive Distributed Stencil Abstraction for GPUs
por: Bhosale, Aditya, et al.
Publicado: (2025) -
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
por: Chang, Li-Wen, et al.
Publicado: (2024)