Saved in:
| Main Authors: | Maeda, Yuki, Taura, Kenjiro |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.05982 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parallel Joinable B-Trees in the Fork-Join I/O Model
by: Goodrich, Michael, et al.
Published: (2025)
by: Goodrich, Michael, et al.
Published: (2025)
GPU-Resident Gaussian Process Regression Leveraging Asynchronous Tasks with HPX
by: Möllmann, Henrik, et al.
Published: (2026)
by: Möllmann, Henrik, et al.
Published: (2026)
Automatic Tracing in Task-Based Runtime Systems
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
by: Ding, Zhimin, et al.
Published: (2024)
by: Ding, Zhimin, et al.
Published: (2024)
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
by: Stepanek, Lukas
Published: (2026)
by: Stepanek, Lukas
Published: (2026)
Runtime-optimized Multi-way Stream Join Operator for Large-scale Streaming data
by: Hu, Jinlong, et al.
Published: (2024)
by: Hu, Jinlong, et al.
Published: (2024)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
GPU-Based Parallel Computing Methods for Medical Photoacoustic Image Reconstruction
by: Yi, Xinyao, et al.
Published: (2024)
by: Yi, Xinyao, et al.
Published: (2024)
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
by: Shan, Baodi, et al.
Published: (2026)
by: Shan, Baodi, et al.
Published: (2026)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
by: Park, Seongyeon, et al.
Published: (2025)
by: Park, Seongyeon, et al.
Published: (2025)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026)
by: Liu, Ruitao, et al.
Published: (2026)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
by: Jiang, Xuan, et al.
Published: (2024)
by: Jiang, Xuan, et al.
Published: (2024)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
by: Liang, Antian, et al.
Published: (2025)
by: Liang, Antian, et al.
Published: (2025)
INSPIRIT: Optimizing Heterogeneous Task Scheduling through Adaptive Priority in Task-based Runtime Systems
by: Wang, Yiqing, et al.
Published: (2024)
by: Wang, Yiqing, et al.
Published: (2024)
Pragma driven shared memory parallelism in Zig by supporting OpenMP loop directives
by: Kacs, David, et al.
Published: (2024)
by: Kacs, David, et al.
Published: (2024)
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
by: Kong, Z. Jonny, et al.
Published: (2025)
by: Kong, Z. Jonny, et al.
Published: (2025)
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
by: Knorr, Fabian, et al.
Published: (2025)
by: Knorr, Fabian, et al.
Published: (2025)
MonadBFT: Fast, Responsive, Fork-Resistant Streamlined Consensus
by: Jalalzai, Mohammad Mussadiq, et al.
Published: (2025)
by: Jalalzai, Mohammad Mussadiq, et al.
Published: (2025)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
by: Latt, Jonas, et al.
Published: (2025)
by: Latt, Jonas, et al.
Published: (2025)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
by: Xia, Bingzheng, et al.
Published: (2025)
by: Xia, Bingzheng, et al.
Published: (2025)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
by: McFarland, Thomas, et al.
Published: (2025)
by: McFarland, Thomas, et al.
Published: (2025)
Radiation Hydrodynamics at Scale: Comparing MPI and Asynchronous Many-Task Runtimes with FleCSI
by: Strack, Alexander, et al.
Published: (2026)
by: Strack, Alexander, et al.
Published: (2026)
CkIO: Parallel File Input for Over-Decomposed Task-Based Systems
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
by: Xia, Mengchun, et al.
Published: (2026)
by: Xia, Mengchun, et al.
Published: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
Neutron particle transport 3D method of characteristic Multi GPU platform Parallel Computing
by: Zhou, Faguo, et al.
Published: (2025)
by: Zhou, Faguo, et al.
Published: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
by: He, Yongchao, et al.
Published: (2025)
by: He, Yongchao, et al.
Published: (2025)
Virtual Garbage Collector (VGC): A Zone-Based Garbage Collection Architecture for Python's Parallel Runtime
by: M, Abdulla
Published: (2025)
by: M, Abdulla
Published: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
CusADi: A GPU Parallelization Framework for Symbolic Expressions and Optimal Control
by: Jeon, Se Hwan, et al.
Published: (2024)
by: Jeon, Se Hwan, et al.
Published: (2024)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
by: Tsolakis, Christos, et al.
Published: (2024)
by: Tsolakis, Christos, et al.
Published: (2024)
EXaCTz: Guaranteed Extremum Graph and Contour Tree Preservation for Distributed- and GPU-Parallel Lossy Compression
by: Li, Yuxiao, et al.
Published: (2026)
by: Li, Yuxiao, et al.
Published: (2026)
Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs
by: Miniskar, Narasinga Rao, et al.
Published: (2026)
by: Miniskar, Narasinga Rao, et al.
Published: (2026)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks
by: Zhou, Sicheng, et al.
Published: (2025)
by: Zhou, Sicheng, et al.
Published: (2025)
Exploring Performance-Productivity Trade-offs in AMT Runtimes: A Task Bench Study of Itoyori, ItoyoriFBC, HPX, and MPI
by: Lahnor, Torben R., et al.
Published: (2026)
by: Lahnor, Torben R., et al.
Published: (2026)
Similar Items
-
Parallel Joinable B-Trees in the Fork-Join I/O Model
by: Goodrich, Michael, et al.
Published: (2025) -
GPU-Resident Gaussian Process Regression Leveraging Asynchronous Tasks with HPX
by: Möllmann, Henrik, et al.
Published: (2026) -
Automatic Tracing in Task-Based Runtime Systems
by: Yadav, Rohan, et al.
Published: (2024) -
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
by: Ding, Zhimin, et al.
Published: (2024) -
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
by: Stepanek, Lukas
Published: (2026)