Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
Fuente:
arXiv
Guardado en:
| Autores principales: | Plesner, Andreas, Sørensen, Hans Henrik Brandenborg, Hauberg, Søren |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
por: Geng, Zipei, et al.
Publicado: (2025)
por: Geng, Zipei, et al.
Publicado: (2025)
Logarithmic-Time Geodesically Convex Decomposition in Programmable Matter
por: Hillebrandt, Henning, et al.
Publicado: (2026)
por: Hillebrandt, Henning, et al.
Publicado: (2026)
Accelerating Maximal Biclique Enumeration on GPUs
por: Hsieh, Chou-Ying, et al.
Publicado: (2024)
por: Hsieh, Chou-Ying, et al.
Publicado: (2024)
Optimizing sDTW for AMD GPUs
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
por: Bhosale, Aditya, et al.
Publicado: (2025)
por: Bhosale, Aditya, et al.
Publicado: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
por: Almasri, Mohammad, et al.
Publicado: (2022)
por: Almasri, Mohammad, et al.
Publicado: (2022)
Accurate Performance Predictors for Edge Computing Applications
por: Giannakopoulos, Panagiotis, et al.
Publicado: (2025)
por: Giannakopoulos, Panagiotis, et al.
Publicado: (2025)
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
por: Pan, Qilong, et al.
Publicado: (2025)
por: Pan, Qilong, et al.
Publicado: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
por: Jangda, Abhinav, et al.
Publicado: (2024)
por: Jangda, Abhinav, et al.
Publicado: (2024)
Optimal Workload Placement on Multi-Instance GPUs
por: Turkkan, Bekir, et al.
Publicado: (2024)
por: Turkkan, Bekir, et al.
Publicado: (2024)
Serving Compound Inference Systems on Datacenter GPUs
por: Devata, Sriram, et al.
Publicado: (2026)
por: Devata, Sriram, et al.
Publicado: (2026)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
por: Zhang, Zeyu, et al.
Publicado: (2025)
por: Zhang, Zeyu, et al.
Publicado: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
por: Brock, Benjamin, et al.
Publicado: (2023)
por: Brock, Benjamin, et al.
Publicado: (2023)
Managing Multi Instance GPUs for High Throughput and Energy Savings
por: Saraha, Abhijeet, et al.
Publicado: (2025)
por: Saraha, Abhijeet, et al.
Publicado: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
por: Ernst, Dominik, et al.
Publicado: (2022)
por: Ernst, Dominik, et al.
Publicado: (2022)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
The Expressive Power of Uniform Population Protocols with Logarithmic Space
por: Czerner, Philipp, et al.
Publicado: (2024)
por: Czerner, Philipp, et al.
Publicado: (2024)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
por: Gao, Wei, et al.
Publicado: (2026)
por: Gao, Wei, et al.
Publicado: (2026)
Anonymized Network Sensing using C++26 std::execution on GPUs
por: Mandulak, Michael, et al.
Publicado: (2025)
por: Mandulak, Michael, et al.
Publicado: (2025)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
por: Li, Shiju, et al.
Publicado: (2025)
por: Li, Shiju, et al.
Publicado: (2025)
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
por: Raj, Suman, et al.
Publicado: (2025)
por: Raj, Suman, et al.
Publicado: (2025)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
por: Ekelund, Jonah, et al.
Publicado: (2025)
por: Ekelund, Jonah, et al.
Publicado: (2025)
Round-optimal $n$-Block Broadcast Schedules in Logarithmic Time
por: Träff, Jesper Larsson
Publicado: (2023)
por: Träff, Jesper Larsson
Publicado: (2023)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
por: Tramm, John, et al.
Publicado: (2024)
por: Tramm, John, et al.
Publicado: (2024)
Accelerating high-order continuum kinetic plasma simulations using multiple GPUs
por: Ho, Andrew, et al.
Publicado: (2024)
por: Ho, Andrew, et al.
Publicado: (2024)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
por: He, Xuan, et al.
Publicado: (2025)
por: He, Xuan, et al.
Publicado: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
por: Dwaraknath, Rajat Vadiraj, et al.
Publicado: (2026)
por: Dwaraknath, Rajat Vadiraj, et al.
Publicado: (2026)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
por: Bellavita, Julian, et al.
Publicado: (2025)
por: Bellavita, Julian, et al.
Publicado: (2025)
DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUs
por: Babaei, Amir Fakhim, et al.
Publicado: (2025)
por: Babaei, Amir Fakhim, et al.
Publicado: (2025)
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
por: Graça, Miguel, et al.
Publicado: (2026)
por: Graça, Miguel, et al.
Publicado: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
por: Zhang, Qijun, et al.
Publicado: (2026)
por: Zhang, Qijun, et al.
Publicado: (2026)
Solving Sequential Greedy Problems Distributedly with Sub-Logarithmic Energy Cost
por: Balliu, Alkida, et al.
Publicado: (2024)
por: Balliu, Alkida, et al.
Publicado: (2024)
The Logarithmic Random Bidding for the Parallel Roulette Wheel Selection with Precise Probabilities
por: Nakano, Koji
Publicado: (2024)
por: Nakano, Koji
Publicado: (2024)
How to Rent GPUs on a Budget
por: Li, Zhouzi, et al.
Publicado: (2024)
por: Li, Zhouzi, et al.
Publicado: (2024)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
por: Hui, Xinning, et al.
Publicado: (2024)
por: Hui, Xinning, et al.
Publicado: (2024)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
por: Arima, Eishi, et al.
Publicado: (2024)
por: Arima, Eishi, et al.
Publicado: (2024)
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
por: Li, Zhouzi, et al.
Publicado: (2026)
por: Li, Zhouzi, et al.
Publicado: (2026)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
por: Ting, Hsu-Tzu, et al.
Publicado: (2025)
por: Ting, Hsu-Tzu, et al.
Publicado: (2025)
An inherently parallel H2-ULV factorization for solving dense linear systems on GPUs
por: Ma, Qianxiang, et al.
Publicado: (2025)
por: Ma, Qianxiang, et al.
Publicado: (2025)
Ejemplares similares
-
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
por: Geng, Zipei, et al.
Publicado: (2025) -
Logarithmic-Time Geodesically Convex Decomposition in Programmable Matter
por: Hillebrandt, Henning, et al.
Publicado: (2026) -
Accelerating Maximal Biclique Enumeration on GPUs
por: Hsieh, Chou-Ying, et al.
Publicado: (2024) -
Optimizing sDTW for AMD GPUs
por: Latta-Lin, Daniel, et al.
Publicado: (2024) -
An Adaptive Distributed Stencil Abstraction for GPUs
por: Bhosale, Aditya, et al.
Publicado: (2025)