Optimizing sDTW for AMD GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Latta-Lin, Daniel, Munoz, Sofia Isadora Padilla |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
by: Tramm, John, et al.
Published: (2024)
by: Tramm, John, et al.
Published: (2024)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
by: Lin, Wei-Chen, et al.
Published: (2024)
by: Lin, Wei-Chen, et al.
Published: (2024)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
by: Lurati, Milo, et al.
Published: (2024)
by: Lurati, Milo, et al.
Published: (2024)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
by: Arima, Eishi, et al.
Published: (2024)
by: Arima, Eishi, et al.
Published: (2024)
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024)
by: Hsieh, Chou-Ying, et al.
Published: (2024)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
by: Almasri, Mohammad, et al.
Published: (2022)
by: Almasri, Mohammad, et al.
Published: (2022)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
by: Liu, Jinyang, et al.
Published: (2023)
by: Liu, Jinyang, et al.
Published: (2023)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
by: Jangda, Abhinav, et al.
Published: (2024)
by: Jangda, Abhinav, et al.
Published: (2024)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
Serving Compound Inference Systems on Datacenter GPUs
by: Devata, Sriram, et al.
Published: (2026)
by: Devata, Sriram, et al.
Published: (2026)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
by: Deshmukh, Aadesh, et al.
Published: (2025)
by: Deshmukh, Aadesh, et al.
Published: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
by: Plesner, Andreas, et al.
Published: (2024)
by: Plesner, Andreas, et al.
Published: (2024)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
exa-AMD: A Scalable Workflow for Accelerating AI-Assisted Materials Discovery and Design
by: Moraru, Maxim, et al.
Published: (2025)
by: Moraru, Maxim, et al.
Published: (2025)
Managing Multi Instance GPUs for High Throughput and Energy Savings
by: Saraha, Abhijeet, et al.
Published: (2025)
by: Saraha, Abhijeet, et al.
Published: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
by: Ernst, Dominik, et al.
Published: (2022)
by: Ernst, Dominik, et al.
Published: (2022)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
by: Lei, Jie, et al.
Published: (2024)
by: Lei, Jie, et al.
Published: (2024)
Anonymized Network Sensing using C++26 std::execution on GPUs
by: Mandulak, Michael, et al.
Published: (2025)
by: Mandulak, Michael, et al.
Published: (2025)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
by: Raj, Suman, et al.
Published: (2025)
by: Raj, Suman, et al.
Published: (2025)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
by: Ekelund, Jonah, et al.
Published: (2025)
by: Ekelund, Jonah, et al.
Published: (2025)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024)
by: Sun, Mo, et al.
Published: (2024)
Accelerating high-order continuum kinetic plasma simulations using multiple GPUs
by: Ho, Andrew, et al.
Published: (2024)
by: Ho, Andrew, et al.
Published: (2024)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
by: Pan, Qilong, et al.
Published: (2025)
by: Pan, Qilong, et al.
Published: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
by: Bellavita, Julian, et al.
Published: (2025)
by: Bellavita, Julian, et al.
Published: (2025)
DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUs
by: Babaei, Amir Fakhim, et al.
Published: (2025)
by: Babaei, Amir Fakhim, et al.
Published: (2025)
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
by: Graça, Miguel, et al.
Published: (2026)
by: Graça, Miguel, et al.
Published: (2026)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
by: Schieffer, Gabin, et al.
Published: (2025)
by: Schieffer, Gabin, et al.
Published: (2025)
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
by: Vanecek, Stepan, et al.
Published: (2025)
by: Vanecek, Stepan, et al.
Published: (2025)
How to Rent GPUs on a Budget
by: Li, Zhouzi, et al.
Published: (2024)
by: Li, Zhouzi, et al.
Published: (2024)
Similar Items
-
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
by: Tramm, John, et al.
Published: (2024) -
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
by: Lin, Wei-Chen, et al.
Published: (2024) -
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026) -
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
by: Lurati, Milo, et al.
Published: (2024) -
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
by: Arima, Eishi, et al.
Published: (2024)