Stream parallel skeleton optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Aldinucci, Marco, Danelutto, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020)
by: Colonnelli, Iacopo, et al.
Published: (2020)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
by: Gallone, Anna, et al.
Published: (2025)
by: Gallone, Anna, et al.
Published: (2025)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
by: Mantha, Pradeep, et al.
Published: (2026)
by: Mantha, Pradeep, et al.
Published: (2026)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
SpaDA: A Spatial Dataflow Architecture Programming Language
by: Gianinazzi, Lukas, et al.
Published: (2025)
by: Gianinazzi, Lukas, et al.
Published: (2025)
A C++17 Thread Pool for High-Performance Scientific Computing
by: Shoshany, Barak
Published: (2021)
by: Shoshany, Barak
Published: (2021)
NVLang: Unified Static Typing for Actor-Based Concurrency on the BEAM
by: Guerreiro, Miguel de Oliveira
Published: (2025)
by: Guerreiro, Miguel de Oliveira
Published: (2025)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
VeriFx: Correct Replicated Data Types for the Masses
by: De Porre, Kevin, et al.
Published: (2022)
by: De Porre, Kevin, et al.
Published: (2022)
PIM-STM: Software Transactional Memory for Processing-In-Memory Systems
by: Lopes, André, et al.
Published: (2024)
by: Lopes, André, et al.
Published: (2024)
TC-GS: A Faster Gaussian Splatting Module Utilizing Tensor Cores
by: Liao, Zimu, et al.
Published: (2025)
by: Liao, Zimu, et al.
Published: (2025)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
by: Singh, Ajay, et al.
Published: (2026)
by: Singh, Ajay, et al.
Published: (2026)
cfdSCOPE: A Fluid-Dynamics Proxy App for Teaching Performance Engineering
by: Arzt, Peter, et al.
Published: (2025)
by: Arzt, Peter, et al.
Published: (2025)
Introducing SWIRL: An Intermediate Representation Language for Scientific Workflows
by: Colonnelli, Iacopo, et al.
Published: (2024)
by: Colonnelli, Iacopo, et al.
Published: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
by: Li, Shigang, et al.
Published: (2020)
by: Li, Shigang, et al.
Published: (2020)
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
by: Hovland, Paul D.
Published: (2025)
by: Hovland, Paul D.
Published: (2025)
Construction of a Byzantine Linearizable SWMR Atomic Register from SWSR Atomic Registers
by: Kshemkalyani, Ajay D., et al.
Published: (2024)
by: Kshemkalyani, Ajay D., et al.
Published: (2024)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
Scalable Concurrent Queues for GPU
by: Shetty, Pratheek Prakash, et al.
Published: (2026)
by: Shetty, Pratheek Prakash, et al.
Published: (2026)
Categorical Message Passing Language (CaMPL) for programmers
by: Hashimoto, Daniel Kiyoshi, et al.
Published: (2026)
by: Hashimoto, Daniel Kiyoshi, et al.
Published: (2026)
PoCL-R: An Open Standard Based Offloading Layer for Heterogeneous Multi-Access Edge Computing with Server Side Scalability
by: Solanti, Jan, et al.
Published: (2023)
by: Solanti, Jan, et al.
Published: (2023)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
CUNQA: a Distributed Quantum Computing emulator for HPC
by: Vázquez-Pérez, Jorge, et al.
Published: (2025)
by: Vázquez-Pérez, Jorge, et al.
Published: (2025)
Accelerating Gravitational $N$-Body Simulations Using the RISC-V-Based Tenstorrent Wormhole
by: Almerol, Jenny Lynn, et al.
Published: (2025)
by: Almerol, Jenny Lynn, et al.
Published: (2025)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
by: Li, Yinghan, et al.
Published: (2025)
by: Li, Yinghan, et al.
Published: (2025)
Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architectures
by: Wu, Kun, et al.
Published: (2023)
by: Wu, Kun, et al.
Published: (2023)
An Evaluation of Massively Parallel Algorithms for DFA Minimization
by: Martens, Jan, et al.
Published: (2024)
by: Martens, Jan, et al.
Published: (2024)
How to Relax Instantly: Elastic Relaxation of Concurrent Data Structures
by: von Geijer, Kåre, et al.
Published: (2024)
by: von Geijer, Kåre, et al.
Published: (2024)
GPU-centric Communication Schemes for HPC and ML Applications
by: Namashivayam, Naveen
Published: (2025)
by: Namashivayam, Naveen
Published: (2025)
The Autonomous Data Language -- Concepts, Design and Formal Verification
by: Franken, Tom T. P., et al.
Published: (2025)
by: Franken, Tom T. P., et al.
Published: (2025)
HierarchicalKV: A GPU Hash Table with Cache Semantics for Continuous Online Embedding Storage
by: Rong, Haidong, et al.
Published: (2026)
by: Rong, Haidong, et al.
Published: (2026)
Similar Items
-
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020) -
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024) -
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
by: Gallone, Anna, et al.
Published: (2025) -
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
by: Mantha, Pradeep, et al.
Published: (2026) -
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)