Gespeichert in:
| Hauptverfasser: | Kashi, Aditya, Koukpaizan, Nicholson, Lu, Hao, Matheson, Michael, Oral, Sarp, Wang, Feiyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.11512 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora
von: Goto, Kazushige, et al.
Veröffentlicht: (2026)
von: Goto, Kazushige, et al.
Veröffentlicht: (2026)
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
von: Mukherjee, Dipayan, et al.
Veröffentlicht: (2024)
von: Mukherjee, Dipayan, et al.
Veröffentlicht: (2024)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024)
von: He, Jiaao, et al.
Veröffentlicht: (2024)
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
von: Kharma, Sami, et al.
Veröffentlicht: (2026)
von: Kharma, Sami, et al.
Veröffentlicht: (2026)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
von: Putra, Jody Almaida
Veröffentlicht: (2026)
von: Putra, Jody Almaida
Veröffentlicht: (2026)
Intent-driven scheduling of backup jobs
von: Dutta, Souvik, et al.
Veröffentlicht: (2024)
von: Dutta, Souvik, et al.
Veröffentlicht: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Serverless Cold Starts and Where to Find Them
von: Joosen, Artjom, et al.
Veröffentlicht: (2024)
von: Joosen, Artjom, et al.
Veröffentlicht: (2024)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
von: Jallouli, Chaimae, et al.
Veröffentlicht: (2026)
von: Jallouli, Chaimae, et al.
Veröffentlicht: (2026)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
von: Yuan, Renzhong, et al.
Veröffentlicht: (2026)
von: Yuan, Renzhong, et al.
Veröffentlicht: (2026)
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
von: Salazar, José Daniel Montoya
Veröffentlicht: (2026)
von: Salazar, José Daniel Montoya
Veröffentlicht: (2026)
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
von: Miletto, Marcelo Cogo, et al.
Veröffentlicht: (2024)
von: Miletto, Marcelo Cogo, et al.
Veröffentlicht: (2024)
Serinv: A Scalable Library for the Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
von: Maillou, Vincent, et al.
Veröffentlicht: (2025)
von: Maillou, Vincent, et al.
Veröffentlicht: (2025)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
von: Liu, Yuyang, et al.
Veröffentlicht: (2025)
von: Liu, Yuyang, et al.
Veröffentlicht: (2025)
Pipit: Scripting the analysis of parallel execution traces
von: Bhatele, Abhinav, et al.
Veröffentlicht: (2023)
von: Bhatele, Abhinav, et al.
Veröffentlicht: (2023)
Automated Programmatic Performance Analysis of Parallel Programs
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
High-Performance Tensor Contraction without Transposition
von: Matthews, Devin A.
Veröffentlicht: (2016)
von: Matthews, Devin A.
Veröffentlicht: (2016)
Efficient Construction of Large Search Spaces for Auto-Tuning
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
von: Kolluru, Saicharan
Veröffentlicht: (2025)
von: Kolluru, Saicharan
Veröffentlicht: (2025)
Code Generation for Near-Roofline Finite Element Actions on GPUs from Symbolic Variational Forms
von: Kulkarni, Kaushik, et al.
Veröffentlicht: (2025)
von: Kulkarni, Kaushik, et al.
Veröffentlicht: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
von: Grbic, Dragana
Veröffentlicht: (2026)
von: Grbic, Dragana
Veröffentlicht: (2026)
Exploiting Spot Instances for Time-Critical Cloud Workloads Using Optimal Randomized Strategies
von: Bhuyan, Neelkamal, et al.
Veröffentlicht: (2026)
von: Bhuyan, Neelkamal, et al.
Veröffentlicht: (2026)
Opportunistic Scheduling for Optimal Spot Instance Savings in the Cloud
von: Bhuyan, Neelkamal, et al.
Veröffentlicht: (2026)
von: Bhuyan, Neelkamal, et al.
Veröffentlicht: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
von: Villalobos, Johansell, et al.
Veröffentlicht: (2025)
GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers
von: Constantinescu, Denisa-Andreea, et al.
Veröffentlicht: (2026)
von: Constantinescu, Denisa-Andreea, et al.
Veröffentlicht: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora
von: Goto, Kazushige, et al.
Veröffentlicht: (2026) -
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
von: Mukherjee, Dipayan, et al.
Veröffentlicht: (2024) -
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024) -
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
von: Kharma, Sami, et al.
Veröffentlicht: (2026) -
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)