DaiFu: In-Situ Crash Recovery for Deep Learning Systems
Fuente:
arXiv
Saved in:
| Main Authors: | He, Zilong, Chen, Pengfei, Zhang, Hongyu, Li, Xiaoyun, Yu, Guangba, Chen, Hongyang, Zheng, Zibin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FaaSRCA: Full Lifecycle Root Cause Analysis for Serverless Applications
by: Huang, Jin, et al.
Published: (2024)
by: Huang, Jin, et al.
Published: (2024)
Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis
by: Huang, Haiyu, et al.
Published: (2024)
by: Huang, Haiyu, et al.
Published: (2024)
A Survey on Failure Analysis and Fault Injection in AI Systems
by: Yu, Guangba, et al.
Published: (2024)
by: Yu, Guangba, et al.
Published: (2024)
InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix Caching
by: Wang, Yilun, et al.
Published: (2025)
by: Wang, Yilun, et al.
Published: (2025)
Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems
by: Wang, Yilun, et al.
Published: (2026)
by: Wang, Yilun, et al.
Published: (2026)
COCA: Generative Root Cause Analysis for Distributed Systems with Code Knowledge
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
by: Yu, Guangba, et al.
Published: (2026)
by: Yu, Guangba, et al.
Published: (2026)
SparseCoder: Identifier-Aware Sparse Transformer for File-Level Code Summarization
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
Tracezip: Efficient Distributed Tracing via Trace Compression
by: Chen, Zhuangbin, et al.
Published: (2025)
by: Chen, Zhuangbin, et al.
Published: (2025)
LogPrism: Unifying Structure and Variable Encoding for Effective Log Compression
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
Trace Sampling 2.0: Code Knowledge Enhanced Span-level Sampling for Distributed Tracing
by: Wu, Yulun, et al.
Published: (2025)
by: Wu, Yulun, et al.
Published: (2025)
An Empirical Study of Interaction Bugs in ROS-based Software
by: Chen, Zhixiang, et al.
Published: (2025)
by: Chen, Zhixiang, et al.
Published: (2025)
iJTyper: An Iterative Type Inference Framework for Java by Integrating Constraint- and Statistically-based Methods
by: Chen, Zhixiang, et al.
Published: (2024)
by: Chen, Zhixiang, et al.
Published: (2024)
Efficiently Detecting Reentrancy Vulnerabilities in Complex Smart Contracts
by: Wang, Zexu, et al.
Published: (2024)
by: Wang, Zexu, et al.
Published: (2024)
Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
by: Kang, Sungmin, et al.
Published: (2025)
by: Kang, Sungmin, et al.
Published: (2025)
Mono2Sls: Automated Monolith-to-Serverless Migration via Multi-Stage Pipeline with Static Analysis
by: Chen, Xingyan, et al.
Published: (2026)
by: Chen, Xingyan, et al.
Published: (2026)
Cast: Automated Resilience Testing for Production Cloud Service Systems
by: Chen, Zhuangbin, et al.
Published: (2026)
by: Chen, Zhuangbin, et al.
Published: (2026)
When to Stop? Towards Efficient Code Generation in LLMs with Excess Token Prevention
by: Guo, Lianghong, et al.
Published: (2024)
by: Guo, Lianghong, et al.
Published: (2024)
Comment Traps: How Defective Commented-out Code Augment Defects in AI-Assisted Code Generation
by: Huang, Yuan, et al.
Published: (2025)
by: Huang, Yuan, et al.
Published: (2025)
WakeMint: Detecting Sleepminting Vulnerabilities in NFT Smart Contracts
by: Xiao, Lei, et al.
Published: (2025)
by: Xiao, Lei, et al.
Published: (2025)
Definition and Detection of Centralization Defects in Smart Contracts
by: Lin, Zewei, et al.
Published: (2024)
by: Lin, Zewei, et al.
Published: (2024)
RLCoder: Reinforcement Learning for Repository-Level Code Completion
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
CrashJS: A NodeJS Benchmark for Automated Crash Reproduction
by: Oliver, Philip, et al.
Published: (2024)
by: Oliver, Philip, et al.
Published: (2024)
Copy-and-Paste? Identifying EVM-Inequivalent Code Smells in Multi-chain Reuse Contracts
by: Wang, Zexu, et al.
Published: (2025)
by: Wang, Zexu, et al.
Published: (2025)
You Augment Me: Exploring ChatGPT-based Data Augmentation for Semantic Code Search
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
Augmenting Smart Contract Decompiler Output through Fine-grained Dependency Analysis and LLM-facilitated Semantic Recovery
by: Liao, Zeqin, et al.
Published: (2025)
by: Liao, Zeqin, et al.
Published: (2025)
An Empirical Study of ChatGPT-Related Projects and Their Issues on GitHub
by: Lin, Zheng, et al.
Published: (2024)
by: Lin, Zheng, et al.
Published: (2024)
KPIRoot+: An Efficient Integrated Framework for Anomaly Detection and Root Cause Analysis in Large-Scale Cloud Systems
by: Gu, Wenwei, et al.
Published: (2025)
by: Gu, Wenwei, et al.
Published: (2025)
Unity is Strength: Enhancing Precision in Reentrancy Vulnerability Detection of Smart Contract Analysis Tools
by: Wang, Zexu, et al.
Published: (2024)
by: Wang, Zexu, et al.
Published: (2024)
SAFE: Harnessing LLM for Scenario-Driven ADS Testing from Multimodal Crash Data
by: Luo, Siwei, et al.
Published: (2025)
by: Luo, Siwei, et al.
Published: (2025)
Cross-Domain Deep Code Search with Meta Learning
by: Chai, Yitian, et al.
Published: (2022)
by: Chai, Yitian, et al.
Published: (2022)
Crash-free Deductive Verifiers
by: Nauta, Wander, et al.
Published: (2026)
by: Nauta, Wander, et al.
Published: (2026)
SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
Crash Report Enhancement with Large Language Models: An Empirical Study
by: Fahim, S M Farah Al, et al.
Published: (2025)
by: Fahim, S M Farah Al, et al.
Published: (2025)
OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution
by: Guo, Lianghong, et al.
Published: (2025)
by: Guo, Lianghong, et al.
Published: (2025)
ConfLogger: Enhance Systems' Configuration Diagnosability through Configuration Logging
by: Shan, Shiwen, et al.
Published: (2025)
by: Shan, Shiwen, et al.
Published: (2025)
CRPWarner: Warning the Risk of Contract-related Rug Pull in DeFi Smart Contracts
by: Lin, Zewei, et al.
Published: (2024)
by: Lin, Zewei, et al.
Published: (2024)
LIDL: LLM Integration Defect Localization via Knowledge Graph-Enhanced Multi-Agent Analysis
by: Tan, Gou, et al.
Published: (2026)
by: Tan, Gou, et al.
Published: (2026)
Similar Items
-
FaaSRCA: Full Lifecycle Root Cause Analysis for Serverless Applications
by: Huang, Jin, et al.
Published: (2024) -
Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis
by: Huang, Haiyu, et al.
Published: (2024) -
A Survey on Failure Analysis and Fault Injection in AI Systems
by: Yu, Guangba, et al.
Published: (2024) -
InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix Caching
by: Wang, Yilun, et al.
Published: (2025) -
Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems
by: Wang, Yilun, et al.
Published: (2026)