Saved in:
| Main Authors: | Plesner, Andreas, Sørensen, Hans Henrik Brandenborg, Hauberg, Søren |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.08729 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
by: Geng, Zipei, et al.
Published: (2025)
by: Geng, Zipei, et al.
Published: (2025)
Logarithmic-Time Geodesically Convex Decomposition in Programmable Matter
by: Hillebrandt, Henning, et al.
Published: (2026)
by: Hillebrandt, Henning, et al.
Published: (2026)
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024)
by: Hsieh, Chou-Ying, et al.
Published: (2024)
Optimizing sDTW for AMD GPUs
by: Latta-Lin, Daniel, et al.
Published: (2024)
by: Latta-Lin, Daniel, et al.
Published: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
by: Almasri, Mohammad, et al.
Published: (2022)
by: Almasri, Mohammad, et al.
Published: (2022)
Accurate Performance Predictors for Edge Computing Applications
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
by: Giannakopoulos, Panagiotis, et al.
Published: (2025)
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
by: Pan, Qilong, et al.
Published: (2025)
by: Pan, Qilong, et al.
Published: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
by: Jangda, Abhinav, et al.
Published: (2024)
by: Jangda, Abhinav, et al.
Published: (2024)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
Serving Compound Inference Systems on Datacenter GPUs
by: Devata, Sriram, et al.
Published: (2026)
by: Devata, Sriram, et al.
Published: (2026)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
Managing Multi Instance GPUs for High Throughput and Energy Savings
by: Saraha, Abhijeet, et al.
Published: (2025)
by: Saraha, Abhijeet, et al.
Published: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
by: Ernst, Dominik, et al.
Published: (2022)
by: Ernst, Dominik, et al.
Published: (2022)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
by: Zhang, Qijun, et al.
Published: (2026)
by: Zhang, Qijun, et al.
Published: (2026)
The Expressive Power of Uniform Population Protocols with Logarithmic Space
by: Czerner, Philipp, et al.
Published: (2024)
by: Czerner, Philipp, et al.
Published: (2024)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
Anonymized Network Sensing using C++26 std::execution on GPUs
by: Mandulak, Michael, et al.
Published: (2025)
by: Mandulak, Michael, et al.
Published: (2025)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
by: Raj, Suman, et al.
Published: (2025)
by: Raj, Suman, et al.
Published: (2025)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
How to Rent GPUs on a Budget
by: Li, Zhouzi, et al.
Published: (2024)
by: Li, Zhouzi, et al.
Published: (2024)
Round-optimal $n$-Block Broadcast Schedules in Logarithmic Time
by: Träff, Jesper Larsson
Published: (2023)
by: Träff, Jesper Larsson
Published: (2023)
Efficient Dynamic MaxFlow Computation on GPUs
by: Kannappan, Shruthi, et al.
Published: (2025)
by: Kannappan, Shruthi, et al.
Published: (2025)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
by: Tramm, John, et al.
Published: (2024)
by: Tramm, John, et al.
Published: (2024)
Accelerating high-order continuum kinetic plasma simulations using multiple GPUs
by: Ho, Andrew, et al.
Published: (2024)
by: Ho, Andrew, et al.
Published: (2024)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
by: Bellavita, Julian, et al.
Published: (2025)
by: Bellavita, Julian, et al.
Published: (2025)
DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUs
by: Babaei, Amir Fakhim, et al.
Published: (2025)
by: Babaei, Amir Fakhim, et al.
Published: (2025)
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
by: Graça, Miguel, et al.
Published: (2026)
by: Graça, Miguel, et al.
Published: (2026)
Solving Sequential Greedy Problems Distributedly with Sub-Logarithmic Energy Cost
by: Balliu, Alkida, et al.
Published: (2024)
by: Balliu, Alkida, et al.
Published: (2024)
The Logarithmic Random Bidding for the Parallel Roulette Wheel Selection with Precise Probabilities
by: Nakano, Koji
Published: (2024)
by: Nakano, Koji
Published: (2024)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
by: Hui, Xinning, et al.
Published: (2024)
by: Hui, Xinning, et al.
Published: (2024)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
by: Arima, Eishi, et al.
Published: (2024)
by: Arima, Eishi, et al.
Published: (2024)
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
by: Li, Zhouzi, et al.
Published: (2026)
by: Li, Zhouzi, et al.
Published: (2026)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
by: Ting, Hsu-Tzu, et al.
Published: (2025)
by: Ting, Hsu-Tzu, et al.
Published: (2025)
Similar Items
-
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
by: Geng, Zipei, et al.
Published: (2025) -
Logarithmic-Time Geodesically Convex Decomposition in Programmable Matter
by: Hillebrandt, Henning, et al.
Published: (2026) -
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024) -
Optimizing sDTW for AMD GPUs
by: Latta-Lin, Daniel, et al.
Published: (2024) -
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)