Wherefore Art Thou? Provenance-Guided Automatic Online Debugging with Lumos
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Jingyuan, Zhang, Lei, Schuermann, Leon, Huang, Gongqi, Netravali, Ravi, Levy, Amit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Cost-Performance Analysis of Cloud-Based Retail Point-of-Sale Systems: A Comparative Study of Google Cloud Platform and Microsoft Azure
di: Pagidoju, Ravi Teja
Pubblicazione: (2026)
di: Pagidoju, Ravi Teja
Pubblicazione: (2026)
Microservices-based Software Systems Reengineering: State-of-the-Art and Future Directions
di: Mohottige, Thakshila Imiya, et al.
Pubblicazione: (2024)
di: Mohottige, Thakshila Imiya, et al.
Pubblicazione: (2024)
A Reference Architecture for Governance of Cloud Native Applications
di: Pourmajidi, William, et al.
Pubblicazione: (2023)
di: Pourmajidi, William, et al.
Pubblicazione: (2023)
Proven Distributed Memory Parallelization of Particle Methods
di: Pahlke, Johannes, et al.
Pubblicazione: (2024)
di: Pahlke, Johannes, et al.
Pubblicazione: (2024)
MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era
di: Zhang, Lei, et al.
Pubblicazione: (2026)
di: Zhang, Lei, et al.
Pubblicazione: (2026)
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
di: Jiang, Zhihan, et al.
Pubblicazione: (2025)
ATOM: Asynchronous Training of Massive Models for Deep Learning in a Decentralized Environment
di: Wu, Xiaofeng, et al.
Pubblicazione: (2024)
di: Wu, Xiaofeng, et al.
Pubblicazione: (2024)
Metronome: Differentiated Delay Scheduling for Serverless Functions
di: Chen, Zhuangbin, et al.
Pubblicazione: (2025)
di: Chen, Zhuangbin, et al.
Pubblicazione: (2025)
Multi-Grained Specifications for Distributed System Model Checking and Verification
di: Ouyang, Lingzhi, et al.
Pubblicazione: (2024)
di: Ouyang, Lingzhi, et al.
Pubblicazione: (2024)
TraceMesh: Scalable and Streaming Sampling for Distributed Traces
di: Chen, Zhuangbin, et al.
Pubblicazione: (2024)
di: Chen, Zhuangbin, et al.
Pubblicazione: (2024)
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
di: Yu, Guangba, et al.
Pubblicazione: (2026)
di: Yu, Guangba, et al.
Pubblicazione: (2026)
MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IR
di: Karchi, Jad El, et al.
Pubblicazione: (2024)
di: Karchi, Jad El, et al.
Pubblicazione: (2024)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
di: Li, Junjie
Pubblicazione: (2024)
di: Li, Junjie
Pubblicazione: (2024)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
di: Wen, Zhongzhen, et al.
Pubblicazione: (2026)
di: Wen, Zhongzhen, et al.
Pubblicazione: (2026)
Supercharging Federated Learning with Flower and NVIDIA FLARE
di: Roth, Holger R., et al.
Pubblicazione: (2024)
di: Roth, Holger R., et al.
Pubblicazione: (2024)
TraceFL: Interpretability-Driven Debugging in Federated Learning via Neuron Provenance
di: Gill, Waris, et al.
Pubblicazione: (2023)
di: Gill, Waris, et al.
Pubblicazione: (2023)
FedDebug: Systematic Debugging for Federated Learning Applications
di: Gill, Waris, et al.
Pubblicazione: (2023)
di: Gill, Waris, et al.
Pubblicazione: (2023)
A Reinforcement Learning Environment for Automatic Code Optimization in the MLIR Compiler
di: Tirichine, Mohammed, et al.
Pubblicazione: (2024)
di: Tirichine, Mohammed, et al.
Pubblicazione: (2024)
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
di: Diehl, Patrick, et al.
Pubblicazione: (2025)
di: Diehl, Patrick, et al.
Pubblicazione: (2025)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
di: Schmid, Larissa, et al.
Pubblicazione: (2024)
di: Schmid, Larissa, et al.
Pubblicazione: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
di: Domke, Jens, et al.
Pubblicazione: (2025)
di: Domke, Jens, et al.
Pubblicazione: (2025)
CloudHeatMap: Heatmap-Based Monitoring for Large-Scale Cloud Systems
di: Sohana, Sarah, et al.
Pubblicazione: (2024)
di: Sohana, Sarah, et al.
Pubblicazione: (2024)
$μ$OpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics
di: Japke, Nils, et al.
Pubblicazione: (2025)
di: Japke, Nils, et al.
Pubblicazione: (2025)
Adaptable TeaStore
di: Bliudze, Simon, et al.
Pubblicazione: (2024)
di: Bliudze, Simon, et al.
Pubblicazione: (2024)
A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPC
di: Sandås, Petter, et al.
Pubblicazione: (2026)
di: Sandås, Petter, et al.
Pubblicazione: (2026)
Building Castles in the Cloud: Architecting Resilient and Scalable Infrastructure
di: Gundla, Naresh Kumar
Pubblicazione: (2024)
di: Gundla, Naresh Kumar
Pubblicazione: (2024)
Histrio: a Serverless Actor System
di: Buttiglieri, Giorgio Natale, et al.
Pubblicazione: (2024)
di: Buttiglieri, Giorgio Natale, et al.
Pubblicazione: (2024)
GitFarm: Git as a Service for Large-Scale Monorepos
di: Dwivedi, Preetam, et al.
Pubblicazione: (2026)
di: Dwivedi, Preetam, et al.
Pubblicazione: (2026)
Predictive Autoscaling for Node.js on Kubernetes: Lower Latency, Right-Sized Capacity
di: Tymoshenko, Ivan, et al.
Pubblicazione: (2026)
di: Tymoshenko, Ivan, et al.
Pubblicazione: (2026)
Efficiently Reproducing Distributed Workflows in Notebook-based Systems
di: Azaz, Talha, et al.
Pubblicazione: (2026)
di: Azaz, Talha, et al.
Pubblicazione: (2026)
AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices
di: Ndadji, Brice Arléon Zemtsop, et al.
Pubblicazione: (2025)
di: Ndadji, Brice Arléon Zemtsop, et al.
Pubblicazione: (2025)
Carbon-aware Software Services
di: Forti, Stefano, et al.
Pubblicazione: (2024)
di: Forti, Stefano, et al.
Pubblicazione: (2024)
CARISMA: CAR-Integrated Service Mesh Architecture
di: Klein, Kevin, et al.
Pubblicazione: (2024)
di: Klein, Kevin, et al.
Pubblicazione: (2024)
Do Large Language Models Understand Performance Optimization?
di: Cui, Bowen, et al.
Pubblicazione: (2025)
di: Cui, Bowen, et al.
Pubblicazione: (2025)
Umbilical Choir: Automated Live Testing for Edge-To-Cloud FaaS Applications
di: Malekabbasi, Mohammadreza, et al.
Pubblicazione: (2025)
di: Malekabbasi, Mohammadreza, et al.
Pubblicazione: (2025)
Specx: a C++ task-based runtime system for heterogeneous distributed architectures
di: Cardosi, Paul, et al.
Pubblicazione: (2023)
di: Cardosi, Paul, et al.
Pubblicazione: (2023)
Container-level Energy Observability in Kubernetes Clusters
di: Pijnacker, Bjorn, et al.
Pubblicazione: (2025)
di: Pijnacker, Bjorn, et al.
Pubblicazione: (2025)
FlowUnits: Extending Dataflow for the Edge-to-Cloud Computing Continuum
di: Chini, Fabio, et al.
Pubblicazione: (2025)
di: Chini, Fabio, et al.
Pubblicazione: (2025)
Learning Recovery Strategies for Dynamic Self-healing in Reactive Systems
di: Sanabria, Mateo, et al.
Pubblicazione: (2024)
di: Sanabria, Mateo, et al.
Pubblicazione: (2024)
SoK: Microservice Architectures from a Dependability Perspective
di: Kažemaks, Dāvis, et al.
Pubblicazione: (2025)
di: Kažemaks, Dāvis, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Cost-Performance Analysis of Cloud-Based Retail Point-of-Sale Systems: A Comparative Study of Google Cloud Platform and Microsoft Azure
di: Pagidoju, Ravi Teja
Pubblicazione: (2026) -
Microservices-based Software Systems Reengineering: State-of-the-Art and Future Directions
di: Mohottige, Thakshila Imiya, et al.
Pubblicazione: (2024) -
A Reference Architecture for Governance of Cloud Native Applications
di: Pourmajidi, William, et al.
Pubblicazione: (2023) -
Proven Distributed Memory Parallelization of Particle Methods
di: Pahlke, Johannes, et al.
Pubblicazione: (2024) -
MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era
di: Zhang, Lei, et al.
Pubblicazione: (2026)