Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
Fuente:
arXiv
Guardado en:
| Autores principales: | Mandarapu, Durga, Nagarajan, Vani, Pelenitsyn, Artem, Kulkarni, Milind |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Staging Blocked Evaluation over Structured Sparse Matrices
por: Das, Pratyush, et al.
Publicado: (2024)
por: Das, Pratyush, et al.
Publicado: (2024)
Rethinking Collision Detection on GPU Ray Tracing Architecture
por: Mandarapu, Durga Keerthi, et al.
Publicado: (2026)
por: Mandarapu, Durga Keerthi, et al.
Publicado: (2026)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
por: Curless, Brian, et al.
Publicado: (2025)
por: Curless, Brian, et al.
Publicado: (2025)
RAFI -- A Ray/Work Forwarding Infrastructure for Data Parallel Multi-Node/Multi-GPU Computing
por: Wald, Ingo, et al.
Publicado: (2026)
por: Wald, Ingo, et al.
Publicado: (2026)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
por: Xia, Yuning, et al.
Publicado: (2026)
por: Xia, Yuning, et al.
Publicado: (2026)
Towards Real-Time Neural Volumetric Rendering on Mobile Devices: A Measurement Study
por: Wang, Zhe, et al.
Publicado: (2024)
por: Wang, Zhe, et al.
Publicado: (2024)
GPU Volume Rendering with Hierarchical Compression Using VDB
por: Zellmann, Stefan, et al.
Publicado: (2025)
por: Zellmann, Stefan, et al.
Publicado: (2025)
Barrier-Augmented Lagrangian for GPU-based Elastodynamic Contact
por: Guo, Dewen, et al.
Publicado: (2024)
por: Guo, Dewen, et al.
Publicado: (2024)
An efficient GPU approach for designing 3D cultural heritage information systems
por: López, Luis, et al.
Publicado: (2025)
por: López, Luis, et al.
Publicado: (2025)
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
por: V., Karthik, et al.
Publicado: (2024)
por: V., Karthik, et al.
Publicado: (2024)
The Energy Cost of Execution-Idle in GPU Clusters
por: Lei, Yiran, et al.
Publicado: (2026)
por: Lei, Yiran, et al.
Publicado: (2026)
Scalable GPU Performance Variability Analysis framework
por: Lahiry, Ankur, et al.
Publicado: (2025)
por: Lahiry, Ankur, et al.
Publicado: (2025)
On the Partitioning of GPU Power among Multi-Instances
por: Vamja, Tirth, et al.
Publicado: (2025)
por: Vamja, Tirth, et al.
Publicado: (2025)
Exploiting ray tracing technology through OptiX to compute particle interactions with cutoff in a 3D environment on GPU
por: David, Algis, et al.
Publicado: (2024)
por: David, Algis, et al.
Publicado: (2024)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025)
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
por: Meneghin, Massimiliano, et al.
Publicado: (2025)
por: Meneghin, Massimiliano, et al.
Publicado: (2025)
Taking GPU Programming Models to Task for Performance Portability
por: Davis, Joshua H., et al.
Publicado: (2024)
por: Davis, Joshua H., et al.
Publicado: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
por: Zhao, Yanbo, et al.
Publicado: (2025)
por: Zhao, Yanbo, et al.
Publicado: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
por: Davis, Joshua H., et al.
Publicado: (2026)
por: Davis, Joshua H., et al.
Publicado: (2026)
THAPI: Tracing Heterogeneous APIs
por: Bekele, Solomon, et al.
Publicado: (2025)
por: Bekele, Solomon, et al.
Publicado: (2025)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
por: Pilliat, Emmanuel
Publicado: (2026)
por: Pilliat, Emmanuel
Publicado: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
por: Islam, Tanzima Z., et al.
Publicado: (2024)
por: Islam, Tanzima Z., et al.
Publicado: (2024)
An Online Probabilistic Distributed Tracing System
por: Toslali, M., et al.
Publicado: (2024)
por: Toslali, M., et al.
Publicado: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
por: Liu, Shifang, et al.
Publicado: (2025)
por: Liu, Shifang, et al.
Publicado: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
por: Jain, Rutwik, et al.
Publicado: (2026)
por: Jain, Rutwik, et al.
Publicado: (2026)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
por: Wang, Yidi, et al.
Publicado: (2024)
por: Wang, Yidi, et al.
Publicado: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
por: Wahlgren, Jacob, et al.
Publicado: (2025)
por: Wahlgren, Jacob, et al.
Publicado: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
por: Oztop, Beste, et al.
Publicado: (2026)
por: Oztop, Beste, et al.
Publicado: (2026)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
por: Wang, Chen, et al.
Publicado: (2025)
por: Wang, Chen, et al.
Publicado: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
por: Villalobos, Johansell, et al.
Publicado: (2025)
por: Villalobos, Johansell, et al.
Publicado: (2025)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
por: Kabamba, Herve Mbikayi, et al.
Publicado: (2023)
por: Kabamba, Herve Mbikayi, et al.
Publicado: (2023)
Vectorization of Gradient Boosting of Decision Trees Prediction in the CatBoost Library for RISC-V Processors
por: Kozinov, Evgeny, et al.
Publicado: (2024)
por: Kozinov, Evgeny, et al.
Publicado: (2024)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
por: Kim, Sukjin, et al.
Publicado: (2025)
por: Kim, Sukjin, et al.
Publicado: (2025)
Operational Strategies for Non-Disruptive Scheduling Transitions in Production HPC Systems
por: MacLachlan, Glen, et al.
Publicado: (2026)
por: MacLachlan, Glen, et al.
Publicado: (2026)
Parallel Ray Tracing of Black Hole Images Using the Schwarzschild Metric
por: Naddell, Liam, et al.
Publicado: (2025)
por: Naddell, Liam, et al.
Publicado: (2025)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
por: He, Jiaao, et al.
Publicado: (2024)
por: He, Jiaao, et al.
Publicado: (2024)
The Landscape of GPU-Centric Communication
por: Unat, Didem, et al.
Publicado: (2024)
por: Unat, Didem, et al.
Publicado: (2024)
Ejemplares similares
-
Staging Blocked Evaluation over Structured Sparse Matrices
por: Das, Pratyush, et al.
Publicado: (2024) -
Rethinking Collision Detection on GPU Ray Tracing Architecture
por: Mandarapu, Durga Keerthi, et al.
Publicado: (2026) -
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
por: Curless, Brian, et al.
Publicado: (2025) -
RAFI -- A Ray/Work Forwarding Infrastructure for Data Parallel Multi-Node/Multi-GPU Computing
por: Wald, Ingo, et al.
Publicado: (2026) -
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
por: Xia, Yuning, et al.
Publicado: (2026)