Saved in:
| Main Authors: | Kashi, Aditya, Koukpaizan, Nicholson, Lu, Hao, Matheson, Michael, Oral, Sarp, Wang, Feiyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.11512 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora
by: Goto, Kazushige, et al.
Published: (2026)
by: Goto, Kazushige, et al.
Published: (2026)
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
by: Mukherjee, Dipayan, et al.
Published: (2024)
by: Mukherjee, Dipayan, et al.
Published: (2024)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)
by: He, Jiaao, et al.
Published: (2024)
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
by: Kharma, Sami, et al.
Published: (2026)
by: Kharma, Sami, et al.
Published: (2026)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
by: Banchelli, Fabio, et al.
Published: (2025)
by: Banchelli, Fabio, et al.
Published: (2025)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
by: Putra, Jody Almaida
Published: (2026)
by: Putra, Jody Almaida
Published: (2026)
Intent-driven scheduling of backup jobs
by: Dutta, Souvik, et al.
Published: (2024)
by: Dutta, Souvik, et al.
Published: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
Serverless Cold Starts and Where to Find Them
by: Joosen, Artjom, et al.
Published: (2024)
by: Joosen, Artjom, et al.
Published: (2024)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
by: Jallouli, Chaimae, et al.
Published: (2026)
by: Jallouli, Chaimae, et al.
Published: (2026)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024)
by: Shen, Siyuan, et al.
Published: (2024)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
by: Salazar, José Daniel Montoya
Published: (2026)
by: Salazar, José Daniel Montoya
Published: (2026)
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
by: Miletto, Marcelo Cogo, et al.
Published: (2024)
by: Miletto, Marcelo Cogo, et al.
Published: (2024)
Serinv: A Scalable Library for the Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
by: Maillou, Vincent, et al.
Published: (2025)
by: Maillou, Vincent, et al.
Published: (2025)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
Profiling and optimization of multi-card GPU machine learning jobs
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
Pipit: Scripting the analysis of parallel execution traces
by: Bhatele, Abhinav, et al.
Published: (2023)
by: Bhatele, Abhinav, et al.
Published: (2023)
Automated Programmatic Performance Analysis of Parallel Programs
by: Cankur, Onur, et al.
Published: (2024)
by: Cankur, Onur, et al.
Published: (2024)
High-Performance Tensor Contraction without Transposition
by: Matthews, Devin A.
Published: (2016)
by: Matthews, Devin A.
Published: (2016)
Efficient Construction of Large Search Spaces for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)
by: Kolluru, Saicharan
Published: (2025)
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
by: Kulkarni, Kaushik, et al.
Published: (2025)
by: Kulkarni, Kaushik, et al.
Published: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
by: Vatsavai, Sairam Sri, et al.
Published: (2025)
by: Vatsavai, Sairam Sri, et al.
Published: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
by: Arif, Moiz, et al.
Published: (2026)
by: Arif, Moiz, et al.
Published: (2026)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
by: Grbic, Dragana
Published: (2026)
by: Grbic, Dragana
Published: (2026)
Exploiting Spot Instances for Time-Critical Cloud Workloads Using Optimal Randomized Strategies
by: Bhuyan, Neelkamal, et al.
Published: (2026)
by: Bhuyan, Neelkamal, et al.
Published: (2026)
Opportunistic Scheduling for Optimal Spot Instance Savings in the Cloud
by: Bhuyan, Neelkamal, et al.
Published: (2026)
by: Bhuyan, Neelkamal, et al.
Published: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
by: Zhuang, Chen, et al.
Published: (2024)
by: Zhuang, Chen, et al.
Published: (2024)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
by: Xu, Jingwei, et al.
Published: (2025)
by: Xu, Jingwei, et al.
Published: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
by: Dutt, Anurag, et al.
Published: (2025)
by: Dutt, Anurag, et al.
Published: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
by: Ganjihal, Sanjeev Rao
Published: (2026)
by: Ganjihal, Sanjeev Rao
Published: (2026)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
by: Suffa, Philipp, et al.
Published: (2024)
by: Suffa, Philipp, et al.
Published: (2024)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
by: Yildiz, Mert, et al.
Published: (2025)
by: Yildiz, Mert, et al.
Published: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
by: Ather, Hammad, et al.
Published: (2024)
by: Ather, Hammad, et al.
Published: (2024)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers
by: Constantinescu, Denisa-Andreea, et al.
Published: (2026)
by: Constantinescu, Denisa-Andreea, et al.
Published: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
Similar Items
-
Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora
by: Goto, Kazushige, et al.
Published: (2026) -
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
by: Mukherjee, Dipayan, et al.
Published: (2024) -
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024) -
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
by: Kharma, Sami, et al.
Published: (2026) -
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
by: Banchelli, Fabio, et al.
Published: (2025)