Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schieffer, Gabin, Shi, Ruimin, Ren, Jie, Peng, Ivy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2024)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
von: Miksits, Samuel, et al.
Veröffentlicht: (2024)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
von: Shi, Ruimin, et al.
Veröffentlicht: (2025)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
Harnessing CUDA-Q's MPS for Tensor Network Simulations of Large-Scale Quantum Circuits
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2025)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
HeteroSTA: A CPU-GPU Heterogeneous Static Timing Analysis Engine with Holistic Industrial Design Support
von: Guo, Zizheng, et al.
Veröffentlicht: (2025)
von: Guo, Zizheng, et al.
Veröffentlicht: (2025)
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
von: Saba, Issa, et al.
Veröffentlicht: (2024)
von: Saba, Issa, et al.
Veröffentlicht: (2024)
Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
von: Shi, Ruimin, et al.
Veröffentlicht: (2026)
OpenCUBE: Building an Open Source Cloud Blueprint with EPI Systems
von: Peng, Ivy, et al.
Veröffentlicht: (2024)
von: Peng, Ivy, et al.
Veröffentlicht: (2024)
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2026)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2026)
A Unified CPU-GPU Protocol for GNN Training
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2024)
Combining GPU and CPU for accelerating evolutionary computing workloads
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
von: Eynaliyev, Rustam, et al.
Veröffentlicht: (2025)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
von: Qiao, Tong, et al.
Veröffentlicht: (2025)
von: Qiao, Tong, et al.
Veröffentlicht: (2025)
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture
von: Yi, Xinyao
Veröffentlicht: (2024)
von: Yi, Xinyao
Veröffentlicht: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
Static Generation of Efficient OpenMP Offload Data Mappings
von: Marzen, Luke, et al.
Veröffentlicht: (2024)
von: Marzen, Luke, et al.
Veröffentlicht: (2024)
Incidence Constraints in Hypergraph Partitioning on GPU
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
von: Ronzani, Marco, et al.
Veröffentlicht: (2026)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
von: Bridges, Patrick G., et al.
Veröffentlicht: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
von: Chen, Jiesong, et al.
Veröffentlicht: (2026)
von: Chen, Jiesong, et al.
Veröffentlicht: (2026)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
von: He, Yongchao, et al.
Veröffentlicht: (2025)
von: He, Yongchao, et al.
Veröffentlicht: (2025)
Distributed OpenMP Offloading of OpenMC on Intel GPU MAX Accelerators
von: Fridman, Yehonatan, et al.
Veröffentlicht: (2024)
von: Fridman, Yehonatan, et al.
Veröffentlicht: (2024)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2024)
Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024) -
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024) -
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025) -
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024) -
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)