LLMPrism: Black-box Performance Diagnosis for Production LLM Training Platforms
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Zhihan, Ren, Rui, Yu, Guangba, Wu, Yulun, Gu, Wenwei, Li, Yichen, Huang, Yujie, Feng, Cong, Yang, Zengyin, Yang, Yongqiang, Lyu, Michael R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
Trace Sampling 2.0: Code Knowledge Enhanced Span-level Sampling for Distributed Tracing
di: Wu, Yulun, et al.
Pubblicazione: (2025)
di: Wu, Yulun, et al.
Pubblicazione: (2025)
Identifying Performance Issues in Cloud Service Systems Based on Relational-Temporal Features
di: Gu, Wenwei, et al.
Pubblicazione: (2023)
di: Gu, Wenwei, et al.
Pubblicazione: (2023)
FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems
di: Huang, Junjie, et al.
Pubblicazione: (2024)
di: Huang, Junjie, et al.
Pubblicazione: (2024)
COCA: Generative Root Cause Analysis for Distributed Systems with Code Knowledge
di: Li, Yichen, et al.
Pubblicazione: (2025)
di: Li, Yichen, et al.
Pubblicazione: (2025)
Demystifying and Extracting Fault-indicating Information from Logs for Failure Diagnosis
di: Huang, Junjie, et al.
Pubblicazione: (2024)
di: Huang, Junjie, et al.
Pubblicazione: (2024)
CCISolver: End-to-End Detection and Repair of Method-Level Code-Comment Inconsistency
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
Larger Is Not Always Better: Exploring Small Open-source Language Models in Logging Statement Generation
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
KPIRoot+: An Efficient Integrated Framework for Anomaly Detection and Root Cause Analysis in Large-Scale Cloud Systems
di: Gu, Wenwei, et al.
Pubblicazione: (2025)
di: Gu, Wenwei, et al.
Pubblicazione: (2025)
End-to-End Automated Logging via Multi-Agent Framework
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
LogUpdater: Automated Detection and Repair of Specific Defects in Logging Statements
di: Zhong, Renyi, et al.
Pubblicazione: (2024)
di: Zhong, Renyi, et al.
Pubblicazione: (2024)
Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities
di: Ma, Yujie, et al.
Pubblicazione: (2026)
di: Ma, Yujie, et al.
Pubblicazione: (2026)
LogPilot: Intent-aware and Scalable Alert Diagnosis for Large-scale Online Service Systems
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems
di: Wang, Yilun, et al.
Pubblicazione: (2026)
di: Wang, Yilun, et al.
Pubblicazione: (2026)
Why Does the LLM Stop Computing: An Empirical Study of User-Reported Failures in Open-Source LLMs
di: Yu, Guangba, et al.
Pubblicazione: (2026)
di: Yu, Guangba, et al.
Pubblicazione: (2026)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
di: Wang, Zirui, et al.
Pubblicazione: (2026)
di: Wang, Zirui, et al.
Pubblicazione: (2026)
Knowledge-aware Alert Aggregation in Large-scale Cloud Systems: a Hybrid Approach
di: Kuang, Jinxi, et al.
Pubblicazione: (2024)
di: Kuang, Jinxi, et al.
Pubblicazione: (2024)
LUNAR: Unsupervised LLM-based Log Parsing
di: Huang, Junjie, et al.
Pubblicazione: (2024)
di: Huang, Junjie, et al.
Pubblicazione: (2024)
MTAD: Tools and Benchmarks for Multivariate Time Series Anomaly Detection
di: Liu, Jinyang, et al.
Pubblicazione: (2024)
di: Liu, Jinyang, et al.
Pubblicazione: (2024)
Single-Language Evidence Is Insufficient for Automated Logging: A Multilingual Benchmark and Empirical Study with LLMs
di: Zhong, Renyi, et al.
Pubblicazione: (2026)
di: Zhong, Renyi, et al.
Pubblicazione: (2026)
HF-DGF: Hybrid Feedback Guided Directed Grey-box Fuzzing
di: Lyu, Guangfa, et al.
Pubblicazione: (2025)
di: Lyu, Guangfa, et al.
Pubblicazione: (2025)
A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
Go Static: Contextualized Logging Statement Generation
di: Li, Yichen, et al.
Pubblicazione: (2024)
di: Li, Yichen, et al.
Pubblicazione: (2024)
LILAC: Log Parsing using LLMs with Adaptive Parsing Cache
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
di: Jiang, Zhihan, et al.
Pubblicazione: (2023)
MicroRes: Versatile Resilience Profiling in Microservices via Degradation Dissemination Indexing
di: Yang, Tianyi, et al.
Pubblicazione: (2022)
di: Yang, Tianyi, et al.
Pubblicazione: (2022)
An Empirical Evaluation of White-box and Black-box Test Case Prioritization Techniques in CPSs Modeled in Simulink
di: Arrieta, Aitor
Pubblicazione: (2025)
di: Arrieta, Aitor
Pubblicazione: (2025)
Exploring the Effectiveness of LLMs in Automated Logging Generation: An Empirical Study
di: Li, Yichen, et al.
Pubblicazione: (2023)
di: Li, Yichen, et al.
Pubblicazione: (2023)
Enhancing LLM-Based Coding Tools through Native Integration of IDE-Derived Static Context
di: Li, Yichen, et al.
Pubblicazione: (2024)
di: Li, Yichen, et al.
Pubblicazione: (2024)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
di: Singh, Jaskirat, et al.
Pubblicazione: (2024)
di: Singh, Jaskirat, et al.
Pubblicazione: (2024)
FaaSRCA: Full Lifecycle Root Cause Analysis for Serverless Applications
di: Huang, Jin, et al.
Pubblicazione: (2024)
di: Huang, Jin, et al.
Pubblicazione: (2024)
UniSage: A Unified and Post-Analysis-Aware Sampling for Microservices
di: Zhu, Zhouruixing, et al.
Pubblicazione: (2025)
di: Zhu, Zhouruixing, et al.
Pubblicazione: (2025)
Smart Hiring Redefined: An Intelligent Recruitment Management Platform
di: Wu, Fangzhe, et al.
Pubblicazione: (2025)
di: Wu, Fangzhe, et al.
Pubblicazione: (2025)
Fast Deterministic Black-box Context-free Grammar Inference
di: Arefin, Mohammad Rifat, et al.
Pubblicazione: (2023)
di: Arefin, Mohammad Rifat, et al.
Pubblicazione: (2023)
A Black-box Testing Framework for Oracle Quantum Programs
di: Long, Peixun, et al.
Pubblicazione: (2025)
di: Long, Peixun, et al.
Pubblicazione: (2025)
DaiFu: In-Situ Crash Recovery for Deep Learning Systems
di: He, Zilong, et al.
Pubblicazione: (2025)
di: He, Zilong, et al.
Pubblicazione: (2025)
SLIM: a Scalable Light-weight Root Cause Analysis for Imbalanced Data in Microservice
di: Ren, Rui, et al.
Pubblicazione: (2024)
di: Ren, Rui, et al.
Pubblicazione: (2024)
Detecting LLM-generated Code with Subtle Modification by Adversarial Training
di: Yin, Xin, et al.
Pubblicazione: (2025)
di: Yin, Xin, et al.
Pubblicazione: (2025)
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
di: Yu, Guangba, et al.
Pubblicazione: (2026)
di: Yu, Guangba, et al.
Pubblicazione: (2026)
Method-level Change-proneness: A Better Metric for Black-box Test Suite Minimization
di: Siam, Md, et al.
Pubblicazione: (2026)
di: Siam, Md, et al.
Pubblicazione: (2026)
LTM: Scalable and Black-box Similarity-based Test Suite Minimization based on Language Models
di: Pan, Rongqi, et al.
Pubblicazione: (2023)
di: Pan, Rongqi, et al.
Pubblicazione: (2023)
Documenti analoghi
-
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
di: Jiang, Zhihan, et al.
Pubblicazione: (2025) -
Trace Sampling 2.0: Code Knowledge Enhanced Span-level Sampling for Distributed Tracing
di: Wu, Yulun, et al.
Pubblicazione: (2025) -
Identifying Performance Issues in Cloud Service Systems Based on Relational-Temporal Features
di: Gu, Wenwei, et al.
Pubblicazione: (2023) -
FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems
di: Huang, Junjie, et al.
Pubblicazione: (2024) -
COCA: Generative Root Cause Analysis for Distributed Systems with Code Knowledge
di: Li, Yichen, et al.
Pubblicazione: (2025)