Saved in:
| Main Author: | Shoshany, Barak |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2105.00613 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024)
by: Aldinucci, Marco, et al.
Published: (2024)
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020)
by: Colonnelli, Iacopo, et al.
Published: (2020)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
by: Singh, Ajay, et al.
Published: (2026)
by: Singh, Ajay, et al.
Published: (2026)
NVLang: Unified Static Typing for Actor-Based Concurrency on the BEAM
by: Guerreiro, Miguel de Oliveira
Published: (2025)
by: Guerreiro, Miguel de Oliveira
Published: (2025)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
by: Mantha, Pradeep, et al.
Published: (2026)
by: Mantha, Pradeep, et al.
Published: (2026)
VeriFx: Correct Replicated Data Types for the Masses
by: De Porre, Kevin, et al.
Published: (2022)
by: De Porre, Kevin, et al.
Published: (2022)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
Introducing SWIRL: An Intermediate Representation Language for Scientific Workflows
by: Colonnelli, Iacopo, et al.
Published: (2024)
by: Colonnelli, Iacopo, et al.
Published: (2024)
SpaDA: A Spatial Dataflow Architecture Programming Language
by: Gianinazzi, Lukas, et al.
Published: (2025)
by: Gianinazzi, Lukas, et al.
Published: (2025)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
CUNQA: a Distributed Quantum Computing emulator for HPC
by: Vázquez-Pérez, Jorge, et al.
Published: (2025)
by: Vázquez-Pérez, Jorge, et al.
Published: (2025)
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
by: Hovland, Paul D.
Published: (2025)
by: Hovland, Paul D.
Published: (2025)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Categorical Message Passing Language (CaMPL) for programmers
by: Hashimoto, Daniel Kiyoshi, et al.
Published: (2026)
by: Hashimoto, Daniel Kiyoshi, et al.
Published: (2026)
How to Relax Instantly: Elastic Relaxation of Concurrent Data Structures
by: von Geijer, Kåre, et al.
Published: (2024)
by: von Geijer, Kåre, et al.
Published: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
by: Gallone, Anna, et al.
Published: (2025)
by: Gallone, Anna, et al.
Published: (2025)
GPU-centric Communication Schemes for HPC and ML Applications
by: Namashivayam, Naveen
Published: (2025)
by: Namashivayam, Naveen
Published: (2025)
An Evaluation of Massively Parallel Algorithms for DFA Minimization
by: Martens, Jan, et al.
Published: (2024)
by: Martens, Jan, et al.
Published: (2024)
Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architectures
by: Wu, Kun, et al.
Published: (2023)
by: Wu, Kun, et al.
Published: (2023)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
Accelerating Gravitational $N$-Body Simulations Using the RISC-V-Based Tenstorrent Wormhole
by: Almerol, Jenny Lynn, et al.
Published: (2025)
by: Almerol, Jenny Lynn, et al.
Published: (2025)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
by: Li, Yinghan, et al.
Published: (2025)
by: Li, Yinghan, et al.
Published: (2025)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
by: Li, Shigang, et al.
Published: (2020)
by: Li, Shigang, et al.
Published: (2020)
A Study of Performance Portability in Plasma Physics Simulations
by: Ruzicka, Josef, et al.
Published: (2024)
by: Ruzicka, Josef, et al.
Published: (2024)
A Formal Semantics of C with OpenMP Parallelism (Extended Version)
by: Du, Ke, et al.
Published: (2026)
by: Du, Ke, et al.
Published: (2026)
Vahana.jl -- A framework (not only) for large-scale agent-based models
by: Fürst, Steffen, et al.
Published: (2024)
by: Fürst, Steffen, et al.
Published: (2024)
Vectorized Adaptive Histograms for Sparse Oblique Forests
by: Lubonja, Ariel, et al.
Published: (2026)
by: Lubonja, Ariel, et al.
Published: (2026)
Construction of a Byzantine Linearizable SWMR Atomic Register from SWSR Atomic Registers
by: Kshemkalyani, Ajay D., et al.
Published: (2024)
by: Kshemkalyani, Ajay D., et al.
Published: (2024)
Agent-based modeling for realistic reproduction of human mobility and contact behavior to evaluate test and isolation strategies in epidemic infectious disease spread
by: Kerkmann, David, et al.
Published: (2024)
by: Kerkmann, David, et al.
Published: (2024)
Space-Fluid Adaptive Sampling by Self-Organisation
by: Casadei, Roberto, et al.
Published: (2022)
by: Casadei, Roberto, et al.
Published: (2022)
Similar Items
-
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022) -
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024) -
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024) -
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025) -
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024)