Leveraging Multi-Instance GPUs through moldable task scheduling
Fuente:
arXiv
Saved in:
| Main Authors: | Villarrubia, Jorge, Costero, Luis, Igual, Francisco D., Olcoz, Katzalin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
by: Papp, Pál András, et al.
Published: (2024)
by: Papp, Pál András, et al.
Published: (2024)
Multiprocessor Scheduling with Memory Constraints: Fundamental Properties and Finding Optimal Solutions
by: Papp, Pál András, et al.
Published: (2025)
by: Papp, Pál András, et al.
Published: (2025)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
by: Villarrubia, Jorge, et al.
Published: (2026)
by: Villarrubia, Jorge, et al.
Published: (2026)
Improving inference time in multi-TPU systems with profiled model segmentation
by: Villarrubia, Jorge, et al.
Published: (2025)
by: Villarrubia, Jorge, et al.
Published: (2025)
Balanced segmentation of CNNs for multi-TPU inference
by: Villarrubia, Jorge, et al.
Published: (2025)
by: Villarrubia, Jorge, et al.
Published: (2025)
Replication in Graph Partitioning and Scheduling Problems
by: Papp, Pál András, et al.
Published: (2026)
by: Papp, Pál András, et al.
Published: (2026)
An Empirical Evaluation of Quantum-Inspired QUBO Methods for Heterogeneous HPC Workflow Mapping and Scheduling
by: Sharma, Aasish Kumar, et al.
Published: (2026)
by: Sharma, Aasish Kumar, et al.
Published: (2026)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
by: Shakeri, Heman, et al.
Published: (2026)
by: Shakeri, Heman, et al.
Published: (2026)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
by: Lendve, Shardul, et al.
Published: (2024)
by: Lendve, Shardul, et al.
Published: (2024)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026)
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026)
Decentralized Optimization in Time-Varying Networks with Arbitrary Delays
by: Ortega, Tomas, et al.
Published: (2024)
by: Ortega, Tomas, et al.
Published: (2024)
cuGenOpt: A GPU-Accelerated General-Purpose Metaheuristic Framework for Combinatorial Optimization
by: Liu, Yuyang
Published: (2026)
by: Liu, Yuyang
Published: (2026)
Compressed and Sparse Models for Non-Convex Decentralized Learning
by: Campbell, Andrew, et al.
Published: (2023)
by: Campbell, Andrew, et al.
Published: (2023)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
by: Chillarón, Mónica, et al.
Published: (2024)
by: Chillarón, Mónica, et al.
Published: (2024)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
by: Siddiqui, Mohammed Humaid, et al.
Published: (2025)
by: Siddiqui, Mohammed Humaid, et al.
Published: (2025)
Efficient parallel implementation of the multiplicative weight update method for graph-based linear programs
by: Ju, Caleb, et al.
Published: (2023)
by: Ju, Caleb, et al.
Published: (2023)
LAMMPS-KOKKOS: Performance Portable Molecular Dynamics Across Exascale Architectures
by: Johansson, Anders, et al.
Published: (2025)
by: Johansson, Anders, et al.
Published: (2025)
Parallel Self-Avoiding Walks for a Low-Autocorrelation Binary Sequences Problem
by: Bošković, Borko, et al.
Published: (2022)
by: Bošković, Borko, et al.
Published: (2022)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
by: Goldman, Amos, et al.
Published: (2026)
by: Goldman, Amos, et al.
Published: (2026)
Speed, power and cost implications for GPU acceleration of Computational Fluid Dynamics on HPC systems
by: Cooper-Baldock, Zachary, et al.
Published: (2024)
by: Cooper-Baldock, Zachary, et al.
Published: (2024)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
CompressedScaffnew: The First Theoretical Double Acceleration of Communication from Local Training and Compression in Distributed Optimization
by: Condat, Laurent, et al.
Published: (2022)
by: Condat, Laurent, et al.
Published: (2022)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
GPU-Initiated Networking for NCCL
by: Hamidouche, Khaled, et al.
Published: (2025)
by: Hamidouche, Khaled, et al.
Published: (2025)
pFedSOP : Accelerating Training Of Personalized Federated Learning Using Second-Order Optimization
by: Sen, Mrinmay, et al.
Published: (2025)
by: Sen, Mrinmay, et al.
Published: (2025)
CRDT-Based Game State Synchronization in Peer-to-Peer VR
by: Dantas, Abel, et al.
Published: (2025)
by: Dantas, Abel, et al.
Published: (2025)
Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
by: Gao, Yunquan, et al.
Published: (2025)
by: Gao, Yunquan, et al.
Published: (2025)
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
by: Lappi, Oskar, et al.
Published: (2025)
by: Lappi, Oskar, et al.
Published: (2025)
Accelerated Training of Federated Learning via Second-Order Methods
by: Sen, Mrinmay, et al.
Published: (2025)
by: Sen, Mrinmay, et al.
Published: (2025)
Communication Compression for Distributed Learning without Control Variates
by: Ortega, Tomas, et al.
Published: (2024)
by: Ortega, Tomas, et al.
Published: (2024)
Communication Compression for Distributed Learning with Aggregate and Server-Guided Feedback
by: Ortega, Tomas, et al.
Published: (2025)
by: Ortega, Tomas, et al.
Published: (2025)
Quantized and Asynchronous Federated Learning
by: Ortega, Tomas, et al.
Published: (2024)
by: Ortega, Tomas, et al.
Published: (2024)
Decentralized Parameter-Free Online Learning with Compressed Gossip
by: Ortega, Tomas, et al.
Published: (2026)
by: Ortega, Tomas, et al.
Published: (2026)
Decentralized Parameter-Free Online Learning
by: Ortega, Tomas, et al.
Published: (2025)
by: Ortega, Tomas, et al.
Published: (2025)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
Similar Items
-
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
by: Papp, Pál András, et al.
Published: (2024) -
Multiprocessor Scheduling with Memory Constraints: Fundamental Properties and Finding Optimal Solutions
by: Papp, Pál András, et al.
Published: (2025) -
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026) -
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
by: Villarrubia, Jorge, et al.
Published: (2026) -
Improving inference time in multi-TPU systems with profiled model segmentation
by: Villarrubia, Jorge, et al.
Published: (2025)