Guardado en:
| Autores principales: | Medel, VÍctor, Arronategui, Unai, Rana, Omer, BaÑares, JosÉ Ángel, Tolosana-Calasanz, Rafael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2402.04491 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Characterising resource management performance in Kubernetes
por: Medel, Víctor, et al.
Publicado: (2024)
por: Medel, Víctor, et al.
Publicado: (2024)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
por: Chu, Xiaoyu, et al.
Publicado: (2025)
por: Chu, Xiaoyu, et al.
Publicado: (2025)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025)
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
por: Tharwani, Jay, et al.
Publicado: (2025)
por: Tharwani, Jay, et al.
Publicado: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
por: Karfakis, George, et al.
Publicado: (2025)
por: Karfakis, George, et al.
Publicado: (2025)
The SAP Cloud Infrastructure Dataset: A Reality Check of Scheduling and Placement of VMs in Cloud Computing
por: Uhlig, Arno, et al.
Publicado: (2025)
por: Uhlig, Arno, et al.
Publicado: (2025)
DREAMS: Decentralized Resource Allocation and Service Management across the Compute Continuum Using Service Affinity
por: Dinh-Tuan, Hai, et al.
Publicado: (2025)
por: Dinh-Tuan, Hai, et al.
Publicado: (2025)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
por: Poptani, Akash, et al.
Publicado: (2025)
por: Poptani, Akash, et al.
Publicado: (2025)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
por: Battaglini-Fischer, Sándor, et al.
Publicado: (2025)
por: Battaglini-Fischer, Sándor, et al.
Publicado: (2025)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
por: Rashid, Md Hasanur, et al.
Publicado: (2026)
por: Rashid, Md Hasanur, et al.
Publicado: (2026)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
por: Mathews, Dhanya R, et al.
Publicado: (2025)
por: Mathews, Dhanya R, et al.
Publicado: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
por: Ather, Hammad, et al.
Publicado: (2024)
por: Ather, Hammad, et al.
Publicado: (2024)
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
por: Landau, Diogo, et al.
Publicado: (2026)
por: Landau, Diogo, et al.
Publicado: (2026)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
por: Rose, Martin, et al.
Publicado: (2025)
por: Rose, Martin, et al.
Publicado: (2025)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
por: Wang, Tuowei, et al.
Publicado: (2024)
por: Wang, Tuowei, et al.
Publicado: (2024)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
por: Andersson, Måns I., et al.
Publicado: (2025)
por: Andersson, Måns I., et al.
Publicado: (2025)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
por: Apanasevich, L., et al.
Publicado: (2024)
por: Apanasevich, L., et al.
Publicado: (2024)
SP-IMPact: A Framework for Static Partitioning Interference Mitigation and Performance Analysis
por: Costa, Diogo, et al.
Publicado: (2025)
por: Costa, Diogo, et al.
Publicado: (2025)
Report on Challenges of Practical Reproducibility for Systems and HPC Computer Science
por: Keahey, Kate, et al.
Publicado: (2025)
por: Keahey, Kate, et al.
Publicado: (2025)
Chopin: An Open Source R-language Tool to Support Spatial Analysis on Parallelizable Infrastructure
por: Song, Insang, et al.
Publicado: (2024)
por: Song, Insang, et al.
Publicado: (2024)
Automated Programmatic Performance Analysis of Parallel Programs
por: Cankur, Onur, et al.
Publicado: (2024)
por: Cankur, Onur, et al.
Publicado: (2024)
Model-driven development of data intensive applications over cloud resources
por: Tolosana-Calasanz, Rafael, et al.
Publicado: (2024)
por: Tolosana-Calasanz, Rafael, et al.
Publicado: (2024)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
por: Cornelius, Melanie, et al.
Publicado: (2025)
por: Cornelius, Melanie, et al.
Publicado: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
por: Zhuang, Chen, et al.
Publicado: (2024)
por: Zhuang, Chen, et al.
Publicado: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
Optimal Parallel Scheduling under Concave Speedup Functions
por: Li, Chengzhang, et al.
Publicado: (2025)
por: Li, Chengzhang, et al.
Publicado: (2025)
WebAssembly and Unikernels: A Comparative Study for Serverless at the Edge
por: Besozzi, Valerio, et al.
Publicado: (2025)
por: Besozzi, Valerio, et al.
Publicado: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
por: Liu, Shifang, et al.
Publicado: (2025)
por: Liu, Shifang, et al.
Publicado: (2025)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
por: Mao, Ying, et al.
Publicado: (2020)
por: Mao, Ying, et al.
Publicado: (2020)
Staging Blocked Evaluation over Structured Sparse Matrices
por: Das, Pratyush, et al.
Publicado: (2024)
por: Das, Pratyush, et al.
Publicado: (2024)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
por: Debnath, Shimul, et al.
Publicado: (2026)
por: Debnath, Shimul, et al.
Publicado: (2026)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
por: Ng, Nathan, et al.
Publicado: (2026)
por: Ng, Nathan, et al.
Publicado: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
por: Lin, Wei-Chen, et al.
Publicado: (2024)
por: Lin, Wei-Chen, et al.
Publicado: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
por: Sun, Tingyang, et al.
Publicado: (2026)
por: Sun, Tingyang, et al.
Publicado: (2026)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
por: Jeffery, Andrew, et al.
Publicado: (2024)
por: Jeffery, Andrew, et al.
Publicado: (2024)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
por: Raffin, Guillaume, et al.
Publicado: (2024)
por: Raffin, Guillaume, et al.
Publicado: (2024)
Bridding OT and PaaS in Edge-to-Cloud Continuum
por: Barrios, Carlos J, et al.
Publicado: (2025)
por: Barrios, Carlos J, et al.
Publicado: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
por: Jain, Rutwik, et al.
Publicado: (2026)
por: Jain, Rutwik, et al.
Publicado: (2026)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
por: Rahimi, Ghazal, et al.
Publicado: (2026)
por: Rahimi, Ghazal, et al.
Publicado: (2026)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
por: Kabamba, Herve Mbikayi, et al.
Publicado: (2023)
por: Kabamba, Herve Mbikayi, et al.
Publicado: (2023)
Ejemplares similares
-
Characterising resource management performance in Kubernetes
por: Medel, Víctor, et al.
Publicado: (2024) -
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
por: Chu, Xiaoyu, et al.
Publicado: (2025) -
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025) -
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
por: Tharwani, Jay, et al.
Publicado: (2025) -
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
por: Karfakis, George, et al.
Publicado: (2025)