FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Battaglini-Fischer, Sándor, Srinivasan, Nishanthi, Szarvas, Bálint László, Chu, Xiaoyu, Iosup, Alexandru |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025)
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
Automated Programmatic Performance Analysis of Parallel Programs
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
von: Cankur, Onur, et al.
Veröffentlicht: (2024)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
PICO: Performance Insights for Collective Operations
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025)
von: Pasqualoni, Saverio, et al.
Veröffentlicht: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
DREAMS: Decentralized Resource Allocation and Service Management across the Compute Continuum Using Service Affinity
von: Dinh-Tuan, Hai, et al.
Veröffentlicht: (2025)
von: Dinh-Tuan, Hai, et al.
Veröffentlicht: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Modeling and Characterizing Service Interference in Dynamic Infrastructures
von: Medel, VÍctor, et al.
Veröffentlicht: (2024)
von: Medel, VÍctor, et al.
Veröffentlicht: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
von: de Morais, Gustavo, et al.
Veröffentlicht: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
von: Lepori, Andrea, et al.
Veröffentlicht: (2025)
von: Lepori, Andrea, et al.
Veröffentlicht: (2025)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
von: Grbic, Dragana
Veröffentlicht: (2026)
von: Grbic, Dragana
Veröffentlicht: (2026)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
A Performance Analysis of BFT Consensus for Blockchains
von: Chan, J. D., et al.
Veröffentlicht: (2024)
von: Chan, J. D., et al.
Veröffentlicht: (2024)
Scalable GPU Performance Variability Analysis framework
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
von: Lahiry, Ankur, et al.
Veröffentlicht: (2025)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024)
von: Qi, S., et al.
Veröffentlicht: (2024)
Performance Debugging through Microarchitectural Sensitivity and Causality Analysis
von: Dutilleul, Alban, et al.
Veröffentlicht: (2024)
von: Dutilleul, Alban, et al.
Veröffentlicht: (2024)
Inductive Loop Analysis for Practical HPC Application Optimization
von: Schaad, Philipp, et al.
Veröffentlicht: (2025)
von: Schaad, Philipp, et al.
Veröffentlicht: (2025)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
von: Bogart, Christopher, et al.
Veröffentlicht: (2025)
von: Bogart, Christopher, et al.
Veröffentlicht: (2025)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
von: León-Vega, Luis G., et al.
Veröffentlicht: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
von: Lurati, Milo, et al.
Veröffentlicht: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2024)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026)
von: Niewenhuis, Dante, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2025) -
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
von: Suman, Shekhar, et al.
Veröffentlicht: (2026) -
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026) -
Automated Programmatic Performance Analysis of Parallel Programs
von: Cankur, Onur, et al.
Veröffentlicht: (2024) -
PASTA: A Modular Program Analysis Tool Framework for Accelerators
von: Lin, Mao, et al.
Veröffentlicht: (2026)