A task-based data-flow methodology for programming heterogeneous systems with multiple accelerator APIs
Fuente:
arXiv
Salvato in:
| Autori principali: | Boné, Aleix, Aguirre, Alejandro, Álvarez, David, Martinez-Ferrer, Pedro J., Beltran, Vicenç |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Thread Scheduling under Oversubscription: A User-Space Framework for Coordinating Multi-runtime and Multi-process Workloads
di: Roca, Aleix, et al.
Pubblicazione: (2026)
di: Roca, Aleix, et al.
Pubblicazione: (2026)
Static task mapping for heterogeneous systems based on series-parallel decompositions
di: Wilhelm, Martin, et al.
Pubblicazione: (2025)
di: Wilhelm, Martin, et al.
Pubblicazione: (2025)
Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
di: Martinez-Ferrer, Pedro J., et al.
Pubblicazione: (2025)
di: Martinez-Ferrer, Pedro J., et al.
Pubblicazione: (2025)
Specx: a C++ task-based runtime system for heterogeneous distributed architectures
di: Cardosi, Paul, et al.
Pubblicazione: (2023)
di: Cardosi, Paul, et al.
Pubblicazione: (2023)
Understanding GPU Triggering APIs for MPI+X Communication
di: Bridges, Patrick G., et al.
Pubblicazione: (2024)
di: Bridges, Patrick G., et al.
Pubblicazione: (2024)
THAPI: Tracing Heterogeneous APIs
di: Bekele, Solomon, et al.
Pubblicazione: (2025)
di: Bekele, Solomon, et al.
Pubblicazione: (2025)
DynoStore: A wide-area distribution system for the management of data over heterogeneous storage
di: Sanchez-Gallegos, Dante D., et al.
Pubblicazione: (2025)
di: Sanchez-Gallegos, Dante D., et al.
Pubblicazione: (2025)
Adaptive Asynchronous Work-Stealing for distributed load-balancing in heterogeneous systems
di: Fernandes, João B., et al.
Pubblicazione: (2024)
di: Fernandes, João B., et al.
Pubblicazione: (2024)
Regent based parallel meshfree LSKUM solver for heterogenous HPC platforms
di: Salil, Sanath, et al.
Pubblicazione: (2024)
di: Salil, Sanath, et al.
Pubblicazione: (2024)
Distributed network for measuring climatic parameters in heterogeneous environments: Application in a greenhouse
di: López-Martínez, Javier, et al.
Pubblicazione: (2024)
di: López-Martínez, Javier, et al.
Pubblicazione: (2024)
HeLoCo: Efficient asynchronous low-communication training under data and device heterogeneity
di: Asif, Abdullah Al, et al.
Pubblicazione: (2026)
di: Asif, Abdullah Al, et al.
Pubblicazione: (2026)
A retrospective on DISPEED -- Leveraging heterogeneity in a drone swarm for IDS execution
di: Lannurien, Vincent, et al.
Pubblicazione: (2025)
di: Lannurien, Vincent, et al.
Pubblicazione: (2025)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
di: Lawenda, Marcin, et al.
Pubblicazione: (2025)
A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates
di: Mao, Zirui, et al.
Pubblicazione: (2023)
di: Mao, Zirui, et al.
Pubblicazione: (2023)
Programming RISC-V accelerators via Fortran
di: Brown, Nick, et al.
Pubblicazione: (2025)
di: Brown, Nick, et al.
Pubblicazione: (2025)
Portable, heterogeneous ensemble workflows at scale using libEnsemble
di: Hudson, Stephen, et al.
Pubblicazione: (2024)
di: Hudson, Stephen, et al.
Pubblicazione: (2024)
Simulating LLM training workloads for heterogeneous compute and network infrastructure
di: Kumar, Sumit, et al.
Pubblicazione: (2025)
di: Kumar, Sumit, et al.
Pubblicazione: (2025)
Energy efficiency optimization of task-parallel codes on asymmetric architectures
di: Costero, Luis, et al.
Pubblicazione: (2024)
di: Costero, Luis, et al.
Pubblicazione: (2024)
Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator
di: Brown, Nick, et al.
Pubblicazione: (2024)
di: Brown, Nick, et al.
Pubblicazione: (2024)
A flexible FPGA accelerator for convolutional neural networks
di: Majumder, Kingshuk, et al.
Pubblicazione: (2019)
di: Majumder, Kingshuk, et al.
Pubblicazione: (2019)
Combining GPU and CPU for accelerating evolutionary computing workloads
di: Eynaliyev, Rustam, et al.
Pubblicazione: (2025)
di: Eynaliyev, Rustam, et al.
Pubblicazione: (2025)
A Survey on Model-heterogeneous Federated Learning: Problems, Methods, and Prospects
di: Fan, Boyu, et al.
Pubblicazione: (2023)
di: Fan, Boyu, et al.
Pubblicazione: (2023)
Dual-pronged deep learning preprocessing on heterogeneous platforms with CPU, Accelerator and CSD
di: Wei, Jia, et al.
Pubblicazione: (2024)
di: Wei, Jia, et al.
Pubblicazione: (2024)
Flotilla: A scalable, modular and resilient federated learning framework for heterogeneous resources
di: Banerjee, Roopkatha, et al.
Pubblicazione: (2025)
di: Banerjee, Roopkatha, et al.
Pubblicazione: (2025)
The integration of heterogeneous resources in the CMS Submission Infrastructure for the LHC Run 3 and beyond
di: Yzquierdo, Antonio Perez-Calero, et al.
Pubblicazione: (2024)
di: Yzquierdo, Antonio Perez-Calero, et al.
Pubblicazione: (2024)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
di: Medeiros, Daniel, et al.
Pubblicazione: (2024)
di: Medeiros, Daniel, et al.
Pubblicazione: (2024)
Performance of a high-order MPI-Kokkos accelerated fluid solver
di: Sporykhin, Filipp, et al.
Pubblicazione: (2025)
di: Sporykhin, Filipp, et al.
Pubblicazione: (2025)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
di: Yi, Xinyao
Pubblicazione: (2024)
di: Yi, Xinyao
Pubblicazione: (2024)
WgPy: GPU-accelerated NumPy-like array library for web browsers
di: Hidaka, Masatoshi, et al.
Pubblicazione: (2025)
di: Hidaka, Masatoshi, et al.
Pubblicazione: (2025)
Non-convex composite federated learning with heterogeneous data
di: Zhang, Jiaojiao, et al.
Pubblicazione: (2025)
di: Zhang, Jiaojiao, et al.
Pubblicazione: (2025)
Collaborative UAVs Multi-task Video Processing Optimization Based on Enhanced Distributed Actor-Critic Networks
di: Rong, Ziqi, et al.
Pubblicazione: (2024)
di: Rong, Ziqi, et al.
Pubblicazione: (2024)
S-VOTE: Similarity-based Voting for Client Selection in Decentralized Federated Learning
di: Sánchez, Pedro Miguel Sánchez, et al.
Pubblicazione: (2025)
di: Sánchez, Pedro Miguel Sánchez, et al.
Pubblicazione: (2025)
Truncated multiplication and batch software SIMD AVX512 implementation for faster Montgomery multiplications and modular exponentiation
di: Didier, Laurent-Stéphane, et al.
Pubblicazione: (2024)
di: Didier, Laurent-Stéphane, et al.
Pubblicazione: (2024)
AdaBridge: Dynamic Data and Computation Reuse for Efficient Multi-task DNN Co-evolution in Edge Systems
di: Wang, Lehao, et al.
Pubblicazione: (2024)
di: Wang, Lehao, et al.
Pubblicazione: (2024)
FedAPTA: Federated Multi-task Learning for Heterogeneous Devices with Adaptive Layer-wise Pruning and Task-aware Aggregation
di: Yu, Zhen, et al.
Pubblicazione: (2025)
di: Yu, Zhen, et al.
Pubblicazione: (2025)
Achieving High-Performance Fault-Tolerant Routing in HyperX Interconnection Networks
di: Camarero, Cristóbal, et al.
Pubblicazione: (2024)
di: Camarero, Cristóbal, et al.
Pubblicazione: (2024)
Resource Allocation in HyperX Networks
di: Cano, Alejandro, et al.
Pubblicazione: (2026)
di: Cano, Alejandro, et al.
Pubblicazione: (2026)
DART: A Solution for Decentralized Federated Learning Model Robustness Analysis
di: Feng, Chao, et al.
Pubblicazione: (2024)
di: Feng, Chao, et al.
Pubblicazione: (2024)
A unified framework to improve the interoperability between HPC and Big Data languages and programming models
di: Piñeiro, César, et al.
Pubblicazione: (2021)
di: Piñeiro, César, et al.
Pubblicazione: (2021)
Recognizing Hereditary Properties in the Presence of Byzantine Nodes
di: Cifuentes-Núñez, David, et al.
Pubblicazione: (2023)
di: Cifuentes-Núñez, David, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Rethinking Thread Scheduling under Oversubscription: A User-Space Framework for Coordinating Multi-runtime and Multi-process Workloads
di: Roca, Aleix, et al.
Pubblicazione: (2026) -
Static task mapping for heterogeneous systems based on series-parallel decompositions
di: Wilhelm, Martin, et al.
Pubblicazione: (2025) -
Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing
di: Martinez-Ferrer, Pedro J., et al.
Pubblicazione: (2025) -
Specx: a C++ task-based runtime system for heterogeneous distributed architectures
di: Cardosi, Paul, et al.
Pubblicazione: (2023) -
Understanding GPU Triggering APIs for MPI+X Communication
di: Bridges, Patrick G., et al.
Pubblicazione: (2024)