FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Junjie, Liu, Jinyang, Chen, Zhuangbin, Jiang, Zhihan, LI, Yichen, Gu, Jiazhen, Feng, Cong, Yang, Zengyin, Yang, Yongqiang, Lyu, Michael R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Demystifying and Extracting Fault-indicating Information from Logs for Failure Diagnosis
di: Huang, Junjie, et al.
Pubblicazione: (2024)
di: Huang, Junjie, et al.
Pubblicazione: (2024)
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
Identifying Performance Issues in Cloud Service Systems Based on Relational-Temporal Features
di: Gu, Wenwei, et al.
Pubblicazione: (2023)
di: Gu, Wenwei, et al.
Pubblicazione: (2023)
Knowledge-aware Alert Aggregation in Large-scale Cloud Systems: a Hybrid Approach
di: Kuang, Jinxi, et al.
Pubblicazione: (2024)
di: Kuang, Jinxi, et al.
Pubblicazione: (2024)
A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
LILAC: Log Parsing using LLMs with Adaptive Parsing Cache
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
KPIRoot+: An Efficient Integrated Framework for Anomaly Detection and Root Cause Analysis in Large-Scale Cloud Systems
di: Gu, Wenwei, et al.
Pubblicazione: (2025)
di: Gu, Wenwei, et al.
Pubblicazione: (2025)
LLMPrism: Black-box Performance Diagnosis for Production LLM Training Platforms
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
LUNAR: Unsupervised LLM-based Log Parsing
di: Huang, Junjie, et al.
Pubblicazione: (2024)
di: Huang, Junjie, et al.
Pubblicazione: (2024)
COCA: Generative Root Cause Analysis for Distributed Systems with Code Knowledge
di: Li, Yichen, et al.
Pubblicazione: (2025)
di: Li, Yichen, et al.
Pubblicazione: (2025)
Go Static: Contextualized Logging Statement Generation
di: Li, Yichen, et al.
Pubblicazione: (2024)
di: Li, Yichen, et al.
Pubblicazione: (2024)
ErrorPrism: Reconstructing Error Propagation Paths in Cloud Service Systems
di: Pu, Junsong, et al.
Pubblicazione: (2025)
di: Pu, Junsong, et al.
Pubblicazione: (2025)
MTAD: Tools and Benchmarks for Multivariate Time Series Anomaly Detection
di: Liu, Jinyang, et al.
Pubblicazione: (2024)
di: Liu, Jinyang, et al.
Pubblicazione: (2024)
MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
di: Jia, Jin, et al.
Pubblicazione: (2026)
di: Jia, Jin, et al.
Pubblicazione: (2026)
LogPilot: Intent-aware and Scalable Alert Diagnosis for Large-scale Online Service Systems
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
TypeScript Repository Indexing for Code Agent Retrieval
di: Pu, Junsong, et al.
Pubblicazione: (2026)
di: Pu, Junsong, et al.
Pubblicazione: (2026)
MicroRacer: Detecting Concurrency Bugs for Cloud Service Systems
di: Deng, Zhiling, et al.
Pubblicazione: (2025)
di: Deng, Zhiling, et al.
Pubblicazione: (2025)
TraceMesh: Scalable and Streaming Sampling for Distributed Traces
di: Chen, Zhuangbin, et al.
Pubblicazione: (2024)
di: Chen, Zhuangbin, et al.
Pubblicazione: (2024)
SBEST: Spectrum-Based Fault Localization Without Fault-Triggering Tests
di: Rafi, Md Nakhla, et al.
Pubblicazione: (2024)
di: Rafi, Md Nakhla, et al.
Pubblicazione: (2024)
MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
MicroRes: Versatile Resilience Profiling in Microservices via Degradation Dissemination Indexing
di: Yang, Tianyi, et al.
Pubblicazione: (2022)
di: Yang, Tianyi, et al.
Pubblicazione: (2022)
Cast: Automated Resilience Testing for Production Cloud Service Systems
di: Chen, Zhuangbin, et al.
Pubblicazione: (2026)
di: Chen, Zhuangbin, et al.
Pubblicazione: (2026)
Neural Fault Injection: Generating Software Faults from Natural Language
di: Cotroneo, Domenico, et al.
Pubblicazione: (2024)
di: Cotroneo, Domenico, et al.
Pubblicazione: (2024)
Large Language Models for Fault Localization: An Empirical Study
di: Xiao, YingJian, et al.
Pubblicazione: (2025)
di: Xiao, YingJian, et al.
Pubblicazione: (2025)
Impact of Large Language Models of Code on Fault Localization
di: Ji, Suhwan, et al.
Pubblicazione: (2024)
di: Ji, Suhwan, et al.
Pubblicazione: (2024)
SieveFL: Hierarchical Runtime-Aware Pruning for Scalable LLM-Based Fault Localization
di: Farzandway, Mahdi, et al.
Pubblicazione: (2026)
di: Farzandway, Mahdi, et al.
Pubblicazione: (2026)
Improved Detection and Diagnosis of Faults in Deep Neural Networks Using Hierarchical and Explainable Classification
di: Jahan, Sigma, et al.
Pubblicazione: (2025)
di: Jahan, Sigma, et al.
Pubblicazione: (2025)
LogPrism: Unifying Structure and Variable Encoding for Effective Log Compression
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
Trace Sampling 2.0: Code Knowledge Enhanced Span-level Sampling for Distributed Tracing
di: Wu, Yulun, et al.
Pubblicazione: (2025)
di: Wu, Yulun, et al.
Pubblicazione: (2025)
Enhancing IR-based Fault Localization using Large Language Models
di: Shao, Shuai, et al.
Pubblicazione: (2024)
di: Shao, Shuai, et al.
Pubblicazione: (2024)
VulKey: Automated Vulnerability Repair Guided by Domain-Specific Repair Patterns
di: Li, Jia, et al.
Pubblicazione: (2026)
di: Li, Jia, et al.
Pubblicazione: (2026)
LLMs-Powered Real-Time Fault Injection: An Approach Toward Intelligent Fault Test Cases Generation
di: Abboush, Mohammad, et al.
Pubblicazione: (2025)
di: Abboush, Mohammad, et al.
Pubblicazione: (2025)
Isolating Compiler Faults via Multiple Pairs of Adversarial Compilation Configurations
di: Li, Qingyang, et al.
Pubblicazione: (2025)
di: Li, Qingyang, et al.
Pubblicazione: (2025)
Explainable Fault Localization for Programming Assignments via LLM-Guided Annotation
di: Liu, Fang, et al.
Pubblicazione: (2025)
di: Liu, Fang, et al.
Pubblicazione: (2025)
TickIt: Leveraging Large Language Models for Automated Ticket Escalation
di: Liu, Fengrui, et al.
Pubblicazione: (2025)
di: Liu, Fengrui, et al.
Pubblicazione: (2025)
TSGuard: Automated User-Centric Incident Diagnosis for AI Workloads in the Cloud
di: Yang, Yitao, et al.
Pubblicazione: (2025)
di: Yang, Yitao, et al.
Pubblicazione: (2025)
Testing for Fault Diversity in Reinforcement Learning
di: Mazouni, Quentin, et al.
Pubblicazione: (2024)
di: Mazouni, Quentin, et al.
Pubblicazione: (2024)
Exploring the Potential and Limitations of Large Language Models for Novice Program Fault Localization
di: Xu, Hexiang, et al.
Pubblicazione: (2025)
di: Xu, Hexiang, et al.
Pubblicazione: (2025)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
di: Jahangirova, Gunel, et al.
Pubblicazione: (2024)
di: Jahangirova, Gunel, et al.
Pubblicazione: (2024)
Tracezip: Efficient Distributed Tracing via Trace Compression
di: Chen, Zhuangbin, et al.
Pubblicazione: (2025)
di: Chen, Zhuangbin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Demystifying and Extracting Fault-indicating Information from Logs for Failure Diagnosis
di: Huang, Junjie, et al.
Pubblicazione: (2024) -
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
di: Jiang, Zhihan, et al.
Pubblicazione: (2025) -
Identifying Performance Issues in Cloud Service Systems Based on Relational-Temporal Features
di: Gu, Wenwei, et al.
Pubblicazione: (2023) -
Knowledge-aware Alert Aggregation in Large-scale Cloud Systems: a Hybrid Approach
di: Kuang, Jinxi, et al.
Pubblicazione: (2024) -
A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)