SIMT/GPU Data Race Verification using ISCC and Intermediary Code Representations: A Case Study
Fuente:
arXiv
Guardado en:
| Autores principales: | Osterhout, Andrew, Gopalakrishnan, Ganesh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HiRace: Accurate and Fast Source-Level Race Checking of GPU Programs
por: Jacobson, John, et al.
Publicado: (2024)
por: Jacobson, John, et al.
Publicado: (2024)
A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates
por: Mao, Zirui, et al.
Publicado: (2023)
por: Mao, Zirui, et al.
Publicado: (2023)
Fast Topology-Aware Lossy Data Compression with Full Preservation of Critical Points and Local Order
por: Fallin, Alex, et al.
Publicado: (2026)
por: Fallin, Alex, et al.
Publicado: (2026)
TopoSZp: Lightweight Topology-Aware Error-controlled Compression for Scientific Data
por: Agarwal, Tripti, et al.
Publicado: (2026)
por: Agarwal, Tripti, et al.
Publicado: (2026)
Characterizing Production GPU Workloads using System-wide Telemetry Data
por: Cankur, Onur, et al.
Publicado: (2025)
por: Cankur, Onur, et al.
Publicado: (2025)
HoSZp: An Efficient Homomorphic Error-bounded Lossy Compressor for Scientific Data
por: Agarwal, Tripti, et al.
Publicado: (2024)
por: Agarwal, Tripti, et al.
Publicado: (2024)
ACC Saturator: Automatic Kernel Optimization for Directive-Based GPU Code
por: Matsumura, Kazuaki, et al.
Publicado: (2023)
por: Matsumura, Kazuaki, et al.
Publicado: (2023)
NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding
por: Wang, Jiamin, et al.
Publicado: (2026)
por: Wang, Jiamin, et al.
Publicado: (2026)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
por: He, Jinghai, et al.
Publicado: (2024)
por: He, Jinghai, et al.
Publicado: (2024)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
por: Siavashi, Ahmad, et al.
Publicado: (2025)
por: Siavashi, Ahmad, et al.
Publicado: (2025)
GPZ: GPU-Accelerated Lossy Compressor for Particle Data
por: Li, Ruoyu, et al.
Publicado: (2025)
por: Li, Ruoyu, et al.
Publicado: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
por: Maurya, Avinash, et al.
Publicado: (2024)
por: Maurya, Avinash, et al.
Publicado: (2024)
Hardware-Aware Reformulation of Convolutions for Efficient Execution on Specialized AI Hardware: A Case Study on NVIDIA Tensor Cores
por: Bikshandi, Ganesh
Publicado: (2026)
por: Bikshandi, Ganesh
Publicado: (2026)
HMTRace: Hardware-Assisted Memory-Tagging based Dynamic Data Race Detection
por: Shastri, Jaidev, et al.
Publicado: (2024)
por: Shastri, Jaidev, et al.
Publicado: (2024)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
por: Yi, Xinyao
Publicado: (2024)
por: Yi, Xinyao
Publicado: (2024)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
por: Schieffer, Gabin, et al.
Publicado: (2024)
por: Schieffer, Gabin, et al.
Publicado: (2024)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
por: Wu, Hao, et al.
Publicado: (2024)
por: Wu, Hao, et al.
Publicado: (2024)
TX-Digital Twin: Visualizing Supercomputer GPU Performance Data Stream
por: Baskakova, Elena, et al.
Publicado: (2026)
por: Baskakova, Elena, et al.
Publicado: (2026)
GPU Sharing with Triples Mode
por: Byun, Chansup, et al.
Publicado: (2024)
por: Byun, Chansup, et al.
Publicado: (2024)
GPU-Accelerated Vecchia Approximations of Gaussian Processes for Geospatial Data using Batched Matrix Computations
por: Pan, Qilong, et al.
Publicado: (2024)
por: Pan, Qilong, et al.
Publicado: (2024)
Intel(R) SHMEM: GPU-initiated OpenSHMEM using SYCL
por: Brooks, Alex, et al.
Publicado: (2024)
por: Brooks, Alex, et al.
Publicado: (2024)
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
por: Latt, Jonas, et al.
Publicado: (2025)
por: Latt, Jonas, et al.
Publicado: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
por: Wang, Tianyu, et al.
Publicado: (2024)
por: Wang, Tianyu, et al.
Publicado: (2024)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
por: Yu, Minchen, et al.
Publicado: (2025)
por: Yu, Minchen, et al.
Publicado: (2025)
Coordinating GPU Data Centers and Power Grid Regulation Service for Exogenous Carbon Benefits
por: Jahanshahi, Ali, et al.
Publicado: (2026)
por: Jahanshahi, Ali, et al.
Publicado: (2026)
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
por: Barrak, Amine, et al.
Publicado: (2025)
por: Barrak, Amine, et al.
Publicado: (2025)
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
por: V., Karthik, et al.
Publicado: (2024)
por: V., Karthik, et al.
Publicado: (2024)
Accelerating Biclique Counting on GPU
por: Qiu, Linshan, et al.
Publicado: (2024)
por: Qiu, Linshan, et al.
Publicado: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
por: Lee, Munkyu, et al.
Publicado: (2024)
por: Lee, Munkyu, et al.
Publicado: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
por: Sojoodi, Amirhossein, et al.
Publicado: (2026)
JAXMg: A multi-GPU linear solver in JAX
por: Wiersema, Roeland
Publicado: (2026)
por: Wiersema, Roeland
Publicado: (2026)
A Unified CPU-GPU Protocol for GNN Training
por: Lin, Yi-Chien, et al.
Publicado: (2024)
por: Lin, Yi-Chien, et al.
Publicado: (2024)
Predictable LLM Serving on GPU Clusters
por: Darzi, Erfan, et al.
Publicado: (2025)
por: Darzi, Erfan, et al.
Publicado: (2025)
DuaLip-GPU Technical Report
por: Dexter, Gregory, et al.
Publicado: (2026)
por: Dexter, Gregory, et al.
Publicado: (2026)
Incidence Constraints in Hypergraph Partitioning on GPU
por: Ronzani, Marco, et al.
Publicado: (2026)
por: Ronzani, Marco, et al.
Publicado: (2026)
GPU Accelerated Sparse Cholesky Factorization
por: Karsavuran, M. Ozan, et al.
Publicado: (2024)
por: Karsavuran, M. Ozan, et al.
Publicado: (2024)
Heat: Satellite's meat is GPU's poison
por: Yuan, Zhehu, et al.
Publicado: (2024)
por: Yuan, Zhehu, et al.
Publicado: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
por: Gu, Jianfeng, et al.
Publicado: (2025)
por: Gu, Jianfeng, et al.
Publicado: (2025)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
por: Fusco, Luigi, et al.
Publicado: (2024)
por: Fusco, Luigi, et al.
Publicado: (2024)
Distributing Context-Aware Shared Memory Data Structures: A Case Study on Singly-Linked Lists
por: Ravishankar, Raaghav, et al.
Publicado: (2024)
por: Ravishankar, Raaghav, et al.
Publicado: (2024)
Ejemplares similares
-
HiRace: Accurate and Fast Source-Level Race Checking of GPU Programs
por: Jacobson, John, et al.
Publicado: (2024) -
A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates
por: Mao, Zirui, et al.
Publicado: (2023) -
Fast Topology-Aware Lossy Data Compression with Full Preservation of Critical Points and Local Order
por: Fallin, Alex, et al.
Publicado: (2026) -
TopoSZp: Lightweight Topology-Aware Error-controlled Compression for Scientific Data
por: Agarwal, Tripti, et al.
Publicado: (2026) -
Characterizing Production GPU Workloads using System-wide Telemetry Data
por: Cankur, Onur, et al.
Publicado: (2025)