Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Curless, Brian, Gowanlock, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
von: Mandarapu, Durga, et al.
Veröffentlicht: (2023)
von: Mandarapu, Durga, et al.
Veröffentlicht: (2023)
Scalable GPU Performance Variability Analysis framework
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
von: Liu, Hang, et al.
Veröffentlicht: (2026)
von: Liu, Hang, et al.
Veröffentlicht: (2026)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
On the Partitioning of GPU Power among Multi-Instances
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
von: Vamja, Tirth, et al.
Veröffentlicht: (2025)
The Energy Cost of Execution-Idle in GPU Clusters
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
Disaggregated Design for GPU-Based Volumetric Data Structures
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
von: Meneghin, Massimiliano, et al.
Veröffentlicht: (2025)
Taking GPU Programming Models to Task for Performance Portability
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
von: Pilliat, Emmanuel
Veröffentlicht: (2026)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
von: Jain, Rutwik, et al.
Veröffentlicht: (2026)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
von: Xia, Yuning, et al.
Veröffentlicht: (2026)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM
von: Ren, Xuanzhengbo, et al.
Veröffentlicht: (2026)
von: Ren, Xuanzhengbo, et al.
Veröffentlicht: (2026)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
mLR: Scalable Laminography Reconstruction based on Memoization
von: Ma, Bin, et al.
Veröffentlicht: (2025)
von: Ma, Bin, et al.
Veröffentlicht: (2025)
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
von: Liu, Hang, et al.
Veröffentlicht: (2025)
von: Liu, Hang, et al.
Veröffentlicht: (2025)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
von: Pennino, Diego, et al.
Veröffentlicht: (2024)
von: Pennino, Diego, et al.
Veröffentlicht: (2024)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
von: Ramesh, Risshab Srinivas
Veröffentlicht: (2024)
von: Ramesh, Risshab Srinivas
Veröffentlicht: (2024)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024)
von: He, Jiaao, et al.
Veröffentlicht: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
von: Laukemann, Jan, et al.
Veröffentlicht: (2023)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
von: Tu, Jiqun, et al.
Veröffentlicht: (2026)
A Performance Analysis of BFT Consensus for Blockchains
von: Chan, J. D., et al.
Veröffentlicht: (2024)
von: Chan, J. D., et al.
Veröffentlicht: (2024)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
The Landscape of GPU-Centric Communication
von: Unat, Didem, et al.
Veröffentlicht: (2024)
von: Unat, Didem, et al.
Veröffentlicht: (2024)
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
von: Laukemann, Jan, et al.
Veröffentlicht: (2024)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
von: Mandarapu, Durga, et al.
Veröffentlicht: (2023) -
Scalable GPU Performance Variability Analysis framework
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025) -
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
von: Liu, Hang, et al.
Veröffentlicht: (2026) -
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025) -
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)