Modular GPU Programming with Typed Perspectives
Fuente:
arXiv
Saved in:
| Main Authors: | Bansal, Manya, Sainati, Daniel, Cutler, Joseph W., Amarasinghe, Saman, Ragan-Kelley, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VeriFx: Correct Replicated Data Types for the Masses
by: De Porre, Kevin, et al.
Published: (2022)
by: De Porre, Kevin, et al.
Published: (2022)
NVLang: Unified Static Typing for Actor-Based Concurrency on the BEAM
by: Guerreiro, Miguel de Oliveira
Published: (2025)
by: Guerreiro, Miguel de Oliveira
Published: (2025)
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
by: Salazar, José Daniel Montoya
Published: (2026)
by: Salazar, José Daniel Montoya
Published: (2026)
Categorical Message Passing Language (CaMPL) for programmers
by: Hashimoto, Daniel Kiyoshi, et al.
Published: (2026)
by: Hashimoto, Daniel Kiyoshi, et al.
Published: (2026)
Send: Objects, History, and Transactions in a Single-Verb Kernel
by: Goes, Christopher
Published: (2026)
by: Goes, Christopher
Published: (2026)
SpaDA: A Spatial Dataflow Architecture Programming Language
by: Gianinazzi, Lukas, et al.
Published: (2025)
by: Gianinazzi, Lukas, et al.
Published: (2025)
PIM-STM: Software Transactional Memory for Processing-In-Memory Systems
by: Lopes, André, et al.
Published: (2024)
by: Lopes, André, et al.
Published: (2024)
Scaling and Load-Balancing Equi-Joins
by: Metwally, Ahmed
Published: (2022)
by: Metwally, Ahmed
Published: (2022)
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024)
by: Aldinucci, Marco, et al.
Published: (2024)
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020)
by: Colonnelli, Iacopo, et al.
Published: (2020)
Inside VOLT: Designing an Open-Source GPU Compiler
by: Jeong, Shinnung, et al.
Published: (2025)
by: Jeong, Shinnung, et al.
Published: (2025)
Introducing SWIRL: An Intermediate Representation Language for Scientific Workflows
by: Colonnelli, Iacopo, et al.
Published: (2024)
by: Colonnelli, Iacopo, et al.
Published: (2024)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
Hardware-Level QoS Enforcement Features: Technologies, Use Cases, and Research Challenges
by: Larsson, Oliver, et al.
Published: (2025)
by: Larsson, Oliver, et al.
Published: (2025)
Rust vs. C for Python Libraries: Evaluating Rust-Compatible Bindings Toolchains
by: Amaral, Isabella Basso do, et al.
Published: (2025)
by: Amaral, Isabella Basso do, et al.
Published: (2025)
Scalable Concurrent Queues for GPU
by: Shetty, Pratheek Prakash, et al.
Published: (2026)
by: Shetty, Pratheek Prakash, et al.
Published: (2026)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
A C++17 Thread Pool for High-Performance Scientific Computing
by: Shoshany, Barak
Published: (2021)
by: Shoshany, Barak
Published: (2021)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architectures
by: Wu, Kun, et al.
Published: (2023)
by: Wu, Kun, et al.
Published: (2023)
A Framework for the Interoperability of Cloud Platforms: Towards FAIR Data in SAFE Environments
by: Grossman, Robert L., et al.
Published: (2022)
by: Grossman, Robert L., et al.
Published: (2022)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
Massively-Parallel Implementation of Inextensible Elastic Rods Using Inter-block GPU Synchronization
by: Korzeniowski, Przemyslaw, et al.
Published: (2025)
by: Korzeniowski, Przemyslaw, et al.
Published: (2025)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
A Treasure Trove of Performance: Analyzing the IO500 Submission Data
by: Kunkel, Julian, et al.
Published: (2026)
by: Kunkel, Julian, et al.
Published: (2026)
A Formal Semantics of C with OpenMP Parallelism (Extended Version)
by: Du, Ke, et al.
Published: (2026)
by: Du, Ke, et al.
Published: (2026)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
by: Singh, Ajay, et al.
Published: (2026)
by: Singh, Ajay, et al.
Published: (2026)
HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML Platforms
by: Van Delm, Josse, et al.
Published: (2024)
by: Van Delm, Josse, et al.
Published: (2024)
Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization
by: Ruggeri, Giuseppe, et al.
Published: (2025)
by: Ruggeri, Giuseppe, et al.
Published: (2025)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
by: Gallone, Anna, et al.
Published: (2025)
by: Gallone, Anna, et al.
Published: (2025)
GPU-centric Communication Schemes for HPC and ML Applications
by: Namashivayam, Naveen
Published: (2025)
by: Namashivayam, Naveen
Published: (2025)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
by: Mantha, Pradeep, et al.
Published: (2026)
by: Mantha, Pradeep, et al.
Published: (2026)
HierarchicalKV: A GPU Hash Table with Cache Semantics for Continuous Online Embedding Storage
by: Rong, Haidong, et al.
Published: (2026)
by: Rong, Haidong, et al.
Published: (2026)
Planetary computing for data-driven environmental policy-making
by: Ferris, Patrick, et al.
Published: (2023)
by: Ferris, Patrick, et al.
Published: (2023)
Offloading tracing for real-time systems using a scalable cloud infrastructure
by: Schmidt, David Jannis, et al.
Published: (2025)
by: Schmidt, David Jannis, et al.
Published: (2025)
HQP: Sensitivity-Aware Hybrid Quantization and Pruning for Ultra-Low-Latency Edge AI Inference
by: Gopalan, Dinesh, et al.
Published: (2026)
by: Gopalan, Dinesh, et al.
Published: (2026)
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
by: Coimbra, Bruno Moreira, et al.
Published: (2025)
by: Coimbra, Bruno Moreira, et al.
Published: (2025)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
Similar Items
-
VeriFx: Correct Replicated Data Types for the Masses
by: De Porre, Kevin, et al.
Published: (2022) -
NVLang: Unified Static Typing for Actor-Based Concurrency on the BEAM
by: Guerreiro, Miguel de Oliveira
Published: (2025) -
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
by: Salazar, José Daniel Montoya
Published: (2026) -
Categorical Message Passing Language (CaMPL) for programmers
by: Hashimoto, Daniel Kiyoshi, et al.
Published: (2026) -
Send: Objects, History, and Transactions in a Single-Verb Kernel
by: Goes, Christopher
Published: (2026)