TraceFL: Interpretability-Driven Debugging in Federated Learning via Neuron Provenance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gill, Waris, Anwar, Ali, Gulzar, Muhammad Ali |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FedDebug: Systematic Debugging for Federated Learning Applications
von: Gill, Waris, et al.
Veröffentlicht: (2023)
von: Gill, Waris, et al.
Veröffentlicht: (2023)
Wherefore Art Thou? Provenance-Guided Automatic Online Debugging with Lumos
von: Chen, Jingyuan, et al.
Veröffentlicht: (2026)
von: Chen, Jingyuan, et al.
Veröffentlicht: (2026)
TraceMesh: Scalable and Streaming Sampling for Distributed Traces
von: Chen, Zhuangbin, et al.
Veröffentlicht: (2024)
von: Chen, Zhuangbin, et al.
Veröffentlicht: (2024)
Supercharging Federated Learning with Flower and NVIDIA FLARE
von: Roth, Holger R., et al.
Veröffentlicht: (2024)
von: Roth, Holger R., et al.
Veröffentlicht: (2024)
MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IR
von: Karchi, Jad El, et al.
Veröffentlicht: (2024)
von: Karchi, Jad El, et al.
Veröffentlicht: (2024)
Towards Secure Management of Edge-Cloud IoT Microservices using Policy as Code
von: Pallewatta, Samodha, et al.
Veröffentlicht: (2024)
von: Pallewatta, Samodha, et al.
Veröffentlicht: (2024)
An SLO Driven and Cost-Aware Autoscaling Framework for Kubernetes
von: Punniyamoorthy, Vinoth, et al.
Veröffentlicht: (2025)
von: Punniyamoorthy, Vinoth, et al.
Veröffentlicht: (2025)
AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices
von: Ndadji, Brice Arléon Zemtsop, et al.
Veröffentlicht: (2025)
von: Ndadji, Brice Arléon Zemtsop, et al.
Veröffentlicht: (2025)
Multi-Objective Load Balancing for Heterogeneous Edge-Based Object Detection Systems
von: Alqahtani, Daghash K., et al.
Veröffentlicht: (2026)
von: Alqahtani, Daghash K., et al.
Veröffentlicht: (2026)
LADs: Leveraging LLMs for AI-Driven DevOps
von: Khan, Ahmad Faraz, et al.
Veröffentlicht: (2025)
von: Khan, Ahmad Faraz, et al.
Veröffentlicht: (2025)
Learning Recovery Strategies for Dynamic Self-healing in Reactive Systems
von: Sanabria, Mateo, et al.
Veröffentlicht: (2024)
von: Sanabria, Mateo, et al.
Veröffentlicht: (2024)
CSnake: Detecting Self-Sustaining Cascading Failure via Causal Stitching of Fault Propagations
von: Qian, Shangshu, et al.
Veröffentlicht: (2025)
von: Qian, Shangshu, et al.
Veröffentlicht: (2025)
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
von: Jiang, Zhihan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhihan, et al.
Veröffentlicht: (2025)
ATOM: Asynchronous Training of Massive Models for Deep Learning in a Decentralized Environment
von: Wu, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Wu, Xiaofeng, et al.
Veröffentlicht: (2024)
Proven Distributed Memory Parallelization of Particle Methods
von: Pahlke, Johannes, et al.
Veröffentlicht: (2024)
von: Pahlke, Johannes, et al.
Veröffentlicht: (2024)
LibProf: A Python Profiler for Improving Cold Start Performance in Serverless Applications
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2024)
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2024)
Federated Learning With L0 Constraint Via Probabilistic Gates For Sparsity
von: Huthasana, Krishna Harsha Kovelakuntla, et al.
Veröffentlicht: (2025)
von: Huthasana, Krishna Harsha Kovelakuntla, et al.
Veröffentlicht: (2025)
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
von: Diehl, Patrick, et al.
Veröffentlicht: (2025)
von: Diehl, Patrick, et al.
Veröffentlicht: (2025)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
von: Schmid, Larissa, et al.
Veröffentlicht: (2024)
von: Schmid, Larissa, et al.
Veröffentlicht: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
von: Domke, Jens, et al.
Veröffentlicht: (2025)
von: Domke, Jens, et al.
Veröffentlicht: (2025)
CloudHeatMap: Heatmap-Based Monitoring for Large-Scale Cloud Systems
von: Sohana, Sarah, et al.
Veröffentlicht: (2024)
von: Sohana, Sarah, et al.
Veröffentlicht: (2024)
$μ$OpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics
von: Japke, Nils, et al.
Veröffentlicht: (2025)
von: Japke, Nils, et al.
Veröffentlicht: (2025)
Adaptable TeaStore
von: Bliudze, Simon, et al.
Veröffentlicht: (2024)
von: Bliudze, Simon, et al.
Veröffentlicht: (2024)
A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPC
von: Sandås, Petter, et al.
Veröffentlicht: (2026)
von: Sandås, Petter, et al.
Veröffentlicht: (2026)
Building Castles in the Cloud: Architecting Resilient and Scalable Infrastructure
von: Gundla, Naresh Kumar
Veröffentlicht: (2024)
von: Gundla, Naresh Kumar
Veröffentlicht: (2024)
Histrio: a Serverless Actor System
von: Buttiglieri, Giorgio Natale, et al.
Veröffentlicht: (2024)
von: Buttiglieri, Giorgio Natale, et al.
Veröffentlicht: (2024)
GitFarm: Git as a Service for Large-Scale Monorepos
von: Dwivedi, Preetam, et al.
Veröffentlicht: (2026)
von: Dwivedi, Preetam, et al.
Veröffentlicht: (2026)
Predictive Autoscaling for Node.js on Kubernetes: Lower Latency, Right-Sized Capacity
von: Tymoshenko, Ivan, et al.
Veröffentlicht: (2026)
von: Tymoshenko, Ivan, et al.
Veröffentlicht: (2026)
Efficiently Reproducing Distributed Workflows in Notebook-based Systems
von: Azaz, Talha, et al.
Veröffentlicht: (2026)
von: Azaz, Talha, et al.
Veröffentlicht: (2026)
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
von: Yu, Guangba, et al.
Veröffentlicht: (2026)
von: Yu, Guangba, et al.
Veröffentlicht: (2026)
Carbon-aware Software Services
von: Forti, Stefano, et al.
Veröffentlicht: (2024)
von: Forti, Stefano, et al.
Veröffentlicht: (2024)
CARISMA: CAR-Integrated Service Mesh Architecture
von: Klein, Kevin, et al.
Veröffentlicht: (2024)
von: Klein, Kevin, et al.
Veröffentlicht: (2024)
Do Large Language Models Understand Performance Optimization?
von: Cui, Bowen, et al.
Veröffentlicht: (2025)
von: Cui, Bowen, et al.
Veröffentlicht: (2025)
Umbilical Choir: Automated Live Testing for Edge-To-Cloud FaaS Applications
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
Specx: a C++ task-based runtime system for heterogeneous distributed architectures
von: Cardosi, Paul, et al.
Veröffentlicht: (2023)
von: Cardosi, Paul, et al.
Veröffentlicht: (2023)
Container-level Energy Observability in Kubernetes Clusters
von: Pijnacker, Bjorn, et al.
Veröffentlicht: (2025)
von: Pijnacker, Bjorn, et al.
Veröffentlicht: (2025)
FlowUnits: Extending Dataflow for the Edge-to-Cloud Computing Continuum
von: Chini, Fabio, et al.
Veröffentlicht: (2025)
von: Chini, Fabio, et al.
Veröffentlicht: (2025)
SoK: Microservice Architectures from a Dependability Perspective
von: Kažemaks, Dāvis, et al.
Veröffentlicht: (2025)
von: Kažemaks, Dāvis, et al.
Veröffentlicht: (2025)
gFaaS: Enabling Generic Functions in Serverless Computing
von: Chadha, Mohak, et al.
Veröffentlicht: (2024)
von: Chadha, Mohak, et al.
Veröffentlicht: (2024)
ShuffleBench: A Benchmark for Large-Scale Data Shuffling Operations with Distributed Stream Processing Frameworks
von: Henning, Sören, et al.
Veröffentlicht: (2024)
von: Henning, Sören, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FedDebug: Systematic Debugging for Federated Learning Applications
von: Gill, Waris, et al.
Veröffentlicht: (2023) -
Wherefore Art Thou? Provenance-Guided Automatic Online Debugging with Lumos
von: Chen, Jingyuan, et al.
Veröffentlicht: (2026) -
TraceMesh: Scalable and Streaming Sampling for Distributed Traces
von: Chen, Zhuangbin, et al.
Veröffentlicht: (2024) -
Supercharging Federated Learning with Flower and NVIDIA FLARE
von: Roth, Holger R., et al.
Veröffentlicht: (2024) -
MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IR
von: Karchi, Jad El, et al.
Veröffentlicht: (2024)