An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chu, Xiaoyu, Talluri, Sacheendra, Lu, Qingxian, Iosup, Alexandru |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
von: Battaglini-Fischer, Sándor, et al.
Veröffentlicht: (2025)
von: Battaglini-Fischer, Sándor, et al.
Veröffentlicht: (2025)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026)
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026)
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services
von: Talluri, Sacheendra, et al.
Veröffentlicht: (2025)
von: Talluri, Sacheendra, et al.
Veröffentlicht: (2025)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024)
Modeling and Characterizing Service Interference in Dynamic Infrastructures
von: Medel, VÍctor, et al.
Veröffentlicht: (2024)
von: Medel, VÍctor, et al.
Veröffentlicht: (2024)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
DREAMS: Decentralized Resource Allocation and Service Management across the Compute Continuum Using Service Affinity
von: Dinh-Tuan, Hai, et al.
Veröffentlicht: (2025)
von: Dinh-Tuan, Hai, et al.
Veröffentlicht: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
von: Poptani, Akash, et al.
Veröffentlicht: (2025)
von: Poptani, Akash, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
von: Yildiz, Mert, et al.
Veröffentlicht: (2025)
LLMPerf: GPU Performance Modeling meets Large Language Models
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
von: Lepori, Andrea, et al.
Veröffentlicht: (2025)
von: Lepori, Andrea, et al.
Veröffentlicht: (2025)
Taking GPU Programming Models to Task for Performance Portability
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
von: Davis, Joshua H., et al.
Veröffentlicht: (2024)
Ridgeline: A 2D Roofline Model for Distributed Systems
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive Control
von: Nguyen, Chanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Chanh, et al.
Veröffentlicht: (2025)
Modeling the Effect of Data Redundancy on Speedup in MLFMA Near-Field Computation
von: Sadeghi, Morteza
Veröffentlicht: (2025)
von: Sadeghi, Morteza
Veröffentlicht: (2025)
Can Large Language Models Predict Parallel Code Performance?
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM
von: Ren, Xuanzhengbo, et al.
Veröffentlicht: (2026)
von: Ren, Xuanzhengbo, et al.
Veröffentlicht: (2026)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Profiling and optimization of multi-card GPU machine learning jobs
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
von: Battaglini-Fischer, Sándor, et al.
Veröffentlicht: (2025) -
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
von: Nicolae, Radu, et al.
Veröffentlicht: (2026) -
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026) -
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
von: Suman, Shekhar, et al.
Veröffentlicht: (2026) -
Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services
von: Talluri, Sacheendra, et al.
Veröffentlicht: (2025)