Speed, power and cost implications for GPU acceleration of Computational Fluid Dynamics on HPC systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Cooper-Baldock, Zachary, Almirall, Brenda Vara, Inthavong, Kiao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
di: He, Jiaao, et al.
Pubblicazione: (2024)
di: He, Jiaao, et al.
Pubblicazione: (2024)
Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
di: Trappen, Tim, et al.
Pubblicazione: (2025)
di: Trappen, Tim, et al.
Pubblicazione: (2025)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
di: Jo, Myeong Jun
Pubblicazione: (2026)
di: Jo, Myeong Jun
Pubblicazione: (2026)
GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers
di: Constantinescu, Denisa-Andreea, et al.
Pubblicazione: (2026)
di: Constantinescu, Denisa-Andreea, et al.
Pubblicazione: (2026)
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
di: Mukherjee, Dipayan, et al.
Pubblicazione: (2024)
di: Mukherjee, Dipayan, et al.
Pubblicazione: (2024)
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
di: Kharma, Sami, et al.
Pubblicazione: (2026)
di: Kharma, Sami, et al.
Pubblicazione: (2026)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
di: Shi, Jinliang, et al.
Pubblicazione: (2025)
di: Shi, Jinliang, et al.
Pubblicazione: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
di: De Sensi, Daniele, et al.
Pubblicazione: (2024)
di: De Sensi, Daniele, et al.
Pubblicazione: (2024)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
di: De Sensi, Daniele, et al.
Pubblicazione: (2025)
di: De Sensi, Daniele, et al.
Pubblicazione: (2025)
GPUnion: Autonomous GPU Sharing on Campus
di: Li, Yufang, et al.
Pubblicazione: (2025)
di: Li, Yufang, et al.
Pubblicazione: (2025)
GPU-centric Communication Schemes for HPC and ML Applications
di: Namashivayam, Naveen
Pubblicazione: (2025)
di: Namashivayam, Naveen
Pubblicazione: (2025)
Lincoln AI Computing Survey (LAICS) and Trends
di: Reuther, Albert, et al.
Pubblicazione: (2025)
di: Reuther, Albert, et al.
Pubblicazione: (2025)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
di: Topcu, Burak, et al.
Pubblicazione: (2026)
di: Topcu, Burak, et al.
Pubblicazione: (2026)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
di: Abraham, Ojima, et al.
Pubblicazione: (2026)
di: Abraham, Ojima, et al.
Pubblicazione: (2026)
Intent-driven scheduling of backup jobs
di: Dutta, Souvik, et al.
Pubblicazione: (2024)
di: Dutta, Souvik, et al.
Pubblicazione: (2024)
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
di: Mitra, Subhadip
Pubblicazione: (2026)
di: Mitra, Subhadip
Pubblicazione: (2026)
Scalability Evaluation of HPC Multi-GPU Training for ECG-based LLMs
di: Mileski, Dimitar, et al.
Pubblicazione: (2025)
di: Mileski, Dimitar, et al.
Pubblicazione: (2025)
Scheduler-Driven Job Atomization
di: Konopa, Michal, et al.
Pubblicazione: (2025)
di: Konopa, Michal, et al.
Pubblicazione: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
di: Konopa, Michal, et al.
Pubblicazione: (2025)
di: Konopa, Michal, et al.
Pubblicazione: (2025)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
di: Xu, Youxuan, et al.
Pubblicazione: (2025)
di: Xu, Youxuan, et al.
Pubblicazione: (2025)
Laminar: A Probe-First Scheduling Paradigm with Deterministic Runtime Survival
di: Chu, Zhengyan
Pubblicazione: (2026)
di: Chu, Zhengyan
Pubblicazione: (2026)
Flex-MIG: Enabling Distributed Execution on MIG
di: Kim, Myeongsu, et al.
Pubblicazione: (2025)
di: Kim, Myeongsu, et al.
Pubblicazione: (2025)
Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
di: Xie, Tianfang
Pubblicazione: (2026)
di: Xie, Tianfang
Pubblicazione: (2026)
Serverless Cold Starts and Where to Find Them
di: Joosen, Artjom, et al.
Pubblicazione: (2024)
di: Joosen, Artjom, et al.
Pubblicazione: (2024)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
di: Jallouli, Chaimae, et al.
Pubblicazione: (2026)
di: Jallouli, Chaimae, et al.
Pubblicazione: (2026)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
di: Putra, Jody Almaida
Pubblicazione: (2026)
di: Putra, Jody Almaida
Pubblicazione: (2026)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
di: Cheng, Long, et al.
Pubblicazione: (2026)
di: Cheng, Long, et al.
Pubblicazione: (2026)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
di: Shakerdargah, Mohammadali, et al.
Pubblicazione: (2024)
di: Shakerdargah, Mohammadali, et al.
Pubblicazione: (2024)
Studying the Effect of Schedule Preemption on Dynamic Task Graph Scheduling
di: Khodabandehlou, Mohammadali, et al.
Pubblicazione: (2026)
di: Khodabandehlou, Mohammadali, et al.
Pubblicazione: (2026)
RACS-SADL: Robust and Understandable Randomized Consensus in the Cloud
di: Tennage, Pasindu, et al.
Pubblicazione: (2024)
di: Tennage, Pasindu, et al.
Pubblicazione: (2024)
Baxos: Backing off for Robust and Efficient Consensus
di: Tennage, Pasindu, et al.
Pubblicazione: (2022)
di: Tennage, Pasindu, et al.
Pubblicazione: (2022)
Scalable Concurrent Queues for GPU
di: Shetty, Pratheek Prakash, et al.
Pubblicazione: (2026)
di: Shetty, Pratheek Prakash, et al.
Pubblicazione: (2026)
Application-Driven Exascale: The JUPITER Benchmark Suite
di: Herten, Andreas, et al.
Pubblicazione: (2024)
di: Herten, Andreas, et al.
Pubblicazione: (2024)
astroCAMP: A Community Benchmark and Co-Design Framework for Sustainable SKA-Scale Radio Imaging
di: Constantinescu, Denisa-Andreea, et al.
Pubblicazione: (2025)
di: Constantinescu, Denisa-Andreea, et al.
Pubblicazione: (2025)
Exploiting Spot Instances for Time-Critical Cloud Workloads Using Optimal Randomized Strategies
di: Bhuyan, Neelkamal, et al.
Pubblicazione: (2026)
di: Bhuyan, Neelkamal, et al.
Pubblicazione: (2026)
Opportunistic Scheduling for Optimal Spot Instance Savings in the Cloud
di: Bhuyan, Neelkamal, et al.
Pubblicazione: (2026)
di: Bhuyan, Neelkamal, et al.
Pubblicazione: (2026)
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
di: Salazar, José Daniel Montoya
Pubblicazione: (2026)
di: Salazar, José Daniel Montoya
Pubblicazione: (2026)
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
di: Miletto, Marcelo Cogo, et al.
Pubblicazione: (2024)
di: Miletto, Marcelo Cogo, et al.
Pubblicazione: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
di: Peng, Hongwu, et al.
Pubblicazione: (2023)
di: Peng, Hongwu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
di: He, Jiaao, et al.
Pubblicazione: (2024) -
Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
di: Trappen, Tim, et al.
Pubblicazione: (2025) -
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
di: Jo, Myeong Jun
Pubblicazione: (2026) -
GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers
di: Constantinescu, Denisa-Andreea, et al.
Pubblicazione: (2026) -
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
di: Mukherjee, Dipayan, et al.
Pubblicazione: (2024)