FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
Fuente:
arXiv
Salvato in:
| Autori principali: | Dwaraknath, Rajat Vadiraj, Kim, Sungyoon, Pilanci, Mert |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
di: Epstein, Elliot L., et al.
Pubblicazione: (2026)
di: Epstein, Elliot L., et al.
Pubblicazione: (2026)
Sampling on Metric Graphs
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2025)
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2025)
Exploring the Landscape of Distributed Graph Sketching
di: Tench, David, et al.
Pubblicazione: (2024)
di: Tench, David, et al.
Pubblicazione: (2024)
SketchGuard: Scaling Byzantine-Robust Decentralized Federated Learning via Sketch-Based Screening
di: Rangwala, Murtaza, et al.
Pubblicazione: (2025)
di: Rangwala, Murtaza, et al.
Pubblicazione: (2025)
GPU-Parallelizable Randomized Sketch-and-Precondition for Linear Regression using Sparse Sign Sketches
di: Chen, Tyler, et al.
Pubblicazione: (2025)
di: Chen, Tyler, et al.
Pubblicazione: (2025)
Communication Lower Bounds and Algorithms for Sketching with Random Dense Matrices
di: Daas, Hussam Al, et al.
Pubblicazione: (2026)
di: Daas, Hussam Al, et al.
Pubblicazione: (2026)
FedNS: A Fast Sketching Newton-Type Algorithm for Federated Learning
di: Li, Jian, et al.
Pubblicazione: (2024)
di: Li, Jian, et al.
Pubblicazione: (2024)
Sketched Gaussian Mechanism for Private Federated Learning
di: Li, Qiaobo, et al.
Pubblicazione: (2025)
di: Li, Qiaobo, et al.
Pubblicazione: (2025)
DiFuseR: A Distributed Sketch-based Influence Maximization Algorithm for GPUs
di: Göktürk, Gökhan, et al.
Pubblicazione: (2024)
di: Göktürk, Gökhan, et al.
Pubblicazione: (2024)
Distributed Recoverable Sketches (Extended Version)
di: Cohen, Diana, et al.
Pubblicazione: (2025)
di: Cohen, Diana, et al.
Pubblicazione: (2025)
Harmonic Decomposition in Data Sketches
di: Wang, Dingyu
Pubblicazione: (2024)
di: Wang, Dingyu
Pubblicazione: (2024)
FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
di: Zhang, Haoyuan, et al.
Pubblicazione: (2025)
di: Zhang, Haoyuan, et al.
Pubblicazione: (2025)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
di: Bellavita, Julian, et al.
Pubblicazione: (2025)
di: Bellavita, Julian, et al.
Pubblicazione: (2025)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
di: Wu, Shixun, et al.
Pubblicazione: (2024)
di: Wu, Shixun, et al.
Pubblicazione: (2024)
Deterministic Lower Bounds for $k$-Edge Connectivity in the Distributed Sketching Model
di: Robinson, Peter, et al.
Pubblicazione: (2025)
di: Robinson, Peter, et al.
Pubblicazione: (2025)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
di: Li, Aiying, et al.
Pubblicazione: (2026)
di: Li, Aiying, et al.
Pubblicazione: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
di: Li, Shiju, et al.
Pubblicazione: (2025)
di: Li, Shiju, et al.
Pubblicazione: (2025)
Spectral Sentinel: Scalable Byzantine-Robust Decentralized Federated Learning via Sketched Random Matrix Theory on Blockchain
di: Mishra, Animesh
Pubblicazione: (2025)
di: Mishra, Animesh
Pubblicazione: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
di: Jangda, Abhinav, et al.
Pubblicazione: (2024)
di: Jangda, Abhinav, et al.
Pubblicazione: (2024)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
di: Brock, Benjamin, et al.
Pubblicazione: (2023)
di: Brock, Benjamin, et al.
Pubblicazione: (2023)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
A High Performance GPU CountSketch Implementation and Its Application to Multisketching and Least Squares Problems
di: Higgins, Andrew J., et al.
Pubblicazione: (2025)
di: Higgins, Andrew J., et al.
Pubblicazione: (2025)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
di: Chang, Li-Wen, et al.
Pubblicazione: (2024)
di: Chang, Li-Wen, et al.
Pubblicazione: (2024)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
di: Yang, Shuo, et al.
Pubblicazione: (2026)
di: Yang, Shuo, et al.
Pubblicazione: (2026)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2025)
di: Wijeratne, Sasindu, et al.
Pubblicazione: (2025)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
di: Zhang, Lixing, et al.
Pubblicazione: (2026)
di: Zhang, Lixing, et al.
Pubblicazione: (2026)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
di: Huang, Ziyu, et al.
Pubblicazione: (2025)
di: Huang, Ziyu, et al.
Pubblicazione: (2025)
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
di: Xie, Xi, et al.
Pubblicazione: (2024)
di: Xie, Xi, et al.
Pubblicazione: (2024)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
di: Shah, Milan, et al.
Pubblicazione: (2026)
di: Shah, Milan, et al.
Pubblicazione: (2026)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
di: Abubaker, Nabil, et al.
Pubblicazione: (2024)
di: Abubaker, Nabil, et al.
Pubblicazione: (2024)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
di: Bellavita, Julian, et al.
Pubblicazione: (2026)
di: Bellavita, Julian, et al.
Pubblicazione: (2026)
An Adaptive Distributed Stencil Abstraction for GPUs
di: Bhosale, Aditya, et al.
Pubblicazione: (2025)
di: Bhosale, Aditya, et al.
Pubblicazione: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
di: Almasri, Mohammad, et al.
Pubblicazione: (2022)
di: Almasri, Mohammad, et al.
Pubblicazione: (2022)
Optimizing sDTW for AMD GPUs
di: Latta-Lin, Daniel, et al.
Pubblicazione: (2024)
di: Latta-Lin, Daniel, et al.
Pubblicazione: (2024)
Sparse Checkpointing for Fast and Reliable MoE Training
di: Gandhi, Swapnil, et al.
Pubblicazione: (2024)
di: Gandhi, Swapnil, et al.
Pubblicazione: (2024)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
Optimal Workload Placement on Multi-Instance GPUs
di: Turkkan, Bekir, et al.
Pubblicazione: (2024)
di: Turkkan, Bekir, et al.
Pubblicazione: (2024)
FlashMoE: Fast Distributed MoE in a Single Kernel
di: Aimuyo, Osayamen Jonathan, et al.
Pubblicazione: (2025)
di: Aimuyo, Osayamen Jonathan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
di: Epstein, Elliot L., et al.
Pubblicazione: (2026) -
Sampling on Metric Graphs
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2025) -
Exploring the Landscape of Distributed Graph Sketching
di: Tench, David, et al.
Pubblicazione: (2024) -
SketchGuard: Scaling Byzantine-Robust Decentralized Federated Learning via Sketch-Based Screening
di: Rangwala, Murtaza, et al.
Pubblicazione: (2025) -
GPU-Parallelizable Randomized Sketch-and-Precondition for Linear Regression using Sparse Sign Sketches
di: Chen, Tyler, et al.
Pubblicazione: (2025)