RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Xi, Luo, Yuebo, Peng, Hongwu, Ding, Caiwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks Training
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
von: Bellavita, Julian, et al.
Veröffentlicht: (2025)
von: Bellavita, Julian, et al.
Veröffentlicht: (2025)
FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal Compression
von: Mechels, Ben, et al.
Veröffentlicht: (2026)
von: Mechels, Ben, et al.
Veröffentlicht: (2026)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
Accelerating Maximal Biclique Enumeration on GPUs
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
Combining Performance and Productivity: Accelerating the Network Sensing Graph Challenge with GPUs and Commodity Data Science Software
von: Samsi, Siddharth, et al.
Veröffentlicht: (2025)
von: Samsi, Siddharth, et al.
Veröffentlicht: (2025)
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025)
von: Li, Shiju, et al.
Veröffentlicht: (2025)
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
Accelerating high-order continuum kinetic plasma simulations using multiple GPUs
von: Ho, Andrew, et al.
Veröffentlicht: (2024)
von: Ho, Andrew, et al.
Veröffentlicht: (2024)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
von: Graça, Miguel, et al.
Veröffentlicht: (2026)
von: Graça, Miguel, et al.
Veröffentlicht: (2026)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2025)
A Lightweight Neural Network for Accelerating Radiative Transfer Modeling in WRF
von: Fredj, Erick, et al.
Veröffentlicht: (2025)
von: Fredj, Erick, et al.
Veröffentlicht: (2025)
Cuckoo-GPU: Accelerating Cuckoo Filters on Modern GPUs
von: Dortmann, Tim, et al.
Veröffentlicht: (2026)
von: Dortmann, Tim, et al.
Veröffentlicht: (2026)
O(K)-Approximation Coflow Scheduling in K-Core Optical Circuit Switching Networks
von: Wang, Xin, et al.
Veröffentlicht: (2026)
von: Wang, Xin, et al.
Veröffentlicht: (2026)
TRUST: Triangle Counting Reloaded on GPUs
von: Pandey, Santosh, et al.
Veröffentlicht: (2021)
von: Pandey, Santosh, et al.
Veröffentlicht: (2021)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
Anonymized Network Sensing using C++26 std::execution on GPUs
von: Mandulak, Michael, et al.
Veröffentlicht: (2025)
von: Mandulak, Michael, et al.
Veröffentlicht: (2025)
Efficient Column-Wise N:M Pruning on RISC-V CPU
von: Chu, Chi-Wei, et al.
Veröffentlicht: (2025)
von: Chu, Chi-Wei, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
Optimizing sDTW for AMD GPUs
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
Optimal Workload Placement on Multi-Instance GPUs
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
von: Turkkan, Bekir, et al.
Veröffentlicht: (2024)
Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
von: Wang, Kewei, et al.
Veröffentlicht: (2025)
von: Wang, Kewei, et al.
Veröffentlicht: (2025)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
Do GPUs Really Need New Tabular File Formats?
von: Luo, Jigao, et al.
Veröffentlicht: (2026)
von: Luo, Jigao, et al.
Veröffentlicht: (2026)
An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
von: Chen, Chuyan, et al.
Veröffentlicht: (2025)
von: Chen, Chuyan, et al.
Veröffentlicht: (2025)
LR-CNN: Lightweight Row-centric Convolutional Neural Network Training for Memory Reduction
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
Stream-K Optimization and Exploration
von: Rackley, Nick, et al.
Veröffentlicht: (2024)
von: Rackley, Nick, et al.
Veröffentlicht: (2024)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
von: Plesner, Andreas, et al.
Veröffentlicht: (2024)
von: Plesner, Andreas, et al.
Veröffentlicht: (2024)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks Training
von: Peng, Hongwu, et al.
Veröffentlicht: (2023) -
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023) -
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
von: Bellavita, Julian, et al.
Veröffentlicht: (2025) -
FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal Compression
von: Mechels, Ben, et al.
Veröffentlicht: (2026) -
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)