How to Rent GPUs on a Budget
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhouzi, Berg, Benjamin, Mukhopadhyay, Arpan, Harchol-Balter, Mor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
Mean field optimal Core Allocation across Malleable jobs
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
von: Rose, Martin, et al.
Veröffentlicht: (2025)
von: Rose, Martin, et al.
Veröffentlicht: (2025)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jianwei, et al.
Veröffentlicht: (2024)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
von: Panova, Elena, et al.
Veröffentlicht: (2022)
von: Panova, Elena, et al.
Veröffentlicht: (2022)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
Performance and scaling of the LFRic weather and climate model on different generations of HPE Cray EX supercomputers
von: Bull, J. Mark, et al.
Veröffentlicht: (2024)
von: Bull, J. Mark, et al.
Veröffentlicht: (2024)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
von: Zhao, Yanbo, et al.
Veröffentlicht: (2025)
An Experimental Study of Different Aggregation Schemes in Semi-Asynchronous Federated Learning
von: Li, Yunbo, et al.
Veröffentlicht: (2024)
von: Li, Yunbo, et al.
Veröffentlicht: (2024)
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
von: Liu, Hang, et al.
Veröffentlicht: (2025)
von: Liu, Hang, et al.
Veröffentlicht: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
mLR: Scalable Laminography Reconstruction based on Memoization
von: Ma, Bin, et al.
Veröffentlicht: (2025)
von: Ma, Bin, et al.
Veröffentlicht: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
von: Chang, Dali, et al.
Veröffentlicht: (2026)
von: Chang, Dali, et al.
Veröffentlicht: (2026)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
von: Heidari, Sina, et al.
Veröffentlicht: (2026)
von: Heidari, Sina, et al.
Veröffentlicht: (2026)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
von: Raffin, Guillaume, et al.
Veröffentlicht: (2024)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
von: Andersson, Måns I., et al.
Veröffentlicht: (2025)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
DiFuseR: A Distributed Sketch-based Influence Maximization Algorithm for GPUs
von: Göktürk, Gökhan, et al.
Veröffentlicht: (2024)
von: Göktürk, Gökhan, et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
von: Li, Zhouzi, et al.
Veröffentlicht: (2026) -
Mean field optimal Core Allocation across Malleable jobs
von: Li, Zhouzi, et al.
Veröffentlicht: (2026) -
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
von: Berg, Benjamin, et al.
Veröffentlicht: (2024) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
von: Afzal, Ayesha, et al.
Veröffentlicht: (2025) -
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)