The workflow motif: a widely-useful performance diagnosis abstraction for distributed applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Abdi, Mania, Desnoyers, Peter, Crovella, Mark, Sambasivan, Raja R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Visualizing Distributed Traces in Aggregate
di: Samanta, Adrita, et al.
Pubblicazione: (2024)
di: Samanta, Adrita, et al.
Pubblicazione: (2024)
FIRED: a fine-grained robust performance diagnosis framework for cloud applications
di: Xin, Ruyue, et al.
Pubblicazione: (2022)
di: Xin, Ruyue, et al.
Pubblicazione: (2022)
Cloud abstractions for AI workloads
di: Canini, Marco, et al.
Pubblicazione: (2025)
di: Canini, Marco, et al.
Pubblicazione: (2025)
emucxl: an emulation framework for CXL-based disaggregated memory applications
di: Gond, Raja, et al.
Pubblicazione: (2024)
di: Gond, Raja, et al.
Pubblicazione: (2024)
Portable, heterogeneous ensemble workflows at scale using libEnsemble
di: Hudson, Stephen, et al.
Pubblicazione: (2024)
di: Hudson, Stephen, et al.
Pubblicazione: (2024)
Local problems in trees across a wide range of distributed models
di: Dhar, Anubhav, et al.
Pubblicazione: (2024)
di: Dhar, Anubhav, et al.
Pubblicazione: (2024)
MaRDIFlow: A CSE workflow framework for abstracting meta-data from FAIR computational experiments
di: Veluvali, Pavan L., et al.
Pubblicazione: (2024)
di: Veluvali, Pavan L., et al.
Pubblicazione: (2024)
Towards cloud-native scientific workflow management
di: Orzechowski, Michal, et al.
Pubblicazione: (2024)
di: Orzechowski, Michal, et al.
Pubblicazione: (2024)
DynoStore: A wide-area distribution system for the management of data over heterogeneous storage
di: Sanchez-Gallegos, Dante D., et al.
Pubblicazione: (2025)
di: Sanchez-Gallegos, Dante D., et al.
Pubblicazione: (2025)
Trace-based, time-resolved analysis of MPI application performance using standard metrics
di: Haldar, Kingshuk
Pubblicazione: (2025)
di: Haldar, Kingshuk
Pubblicazione: (2025)
Dflow, a Python framework for constructing cloud-native AI-for-Science workflows
di: Liu, Xinzijian, et al.
Pubblicazione: (2024)
di: Liu, Xinzijian, et al.
Pubblicazione: (2024)
An overview of the efficiency and censorship-resistance guarantees of widely-used consensus protocols
di: Alpos, Orestis, et al.
Pubblicazione: (2025)
di: Alpos, Orestis, et al.
Pubblicazione: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
di: Cankur, Onur, et al.
Pubblicazione: (2025)
di: Cankur, Onur, et al.
Pubblicazione: (2025)
Misconfiguration prevention and error cause detection for distributed-cloud applications
di: Ranković, Tamara, et al.
Pubblicazione: (2024)
di: Ranković, Tamara, et al.
Pubblicazione: (2024)
SchEdge: A Dynamic, Multi-agent, and Scalable Scheduling Simulator for IoT Edge
di: Hamedi, Ali, et al.
Pubblicazione: (2025)
di: Hamedi, Ali, et al.
Pubblicazione: (2025)
Holistic generational offsets: Fostering a primitive online abstraction for human vs. machine cognition
di: D'Souza, Shaun, et al.
Pubblicazione: (2018)
di: D'Souza, Shaun, et al.
Pubblicazione: (2018)
Blockchain Epidemic Consensus for Large-Scale Networks
di: Abdi, Siamak, et al.
Pubblicazione: (2025)
di: Abdi, Siamak, et al.
Pubblicazione: (2025)
Estimating CO$_2$ emissions of distributed applications and platforms with SimGrid/Batsim
di: Saraiva, Gabriella, et al.
Pubblicazione: (2025)
di: Saraiva, Gabriella, et al.
Pubblicazione: (2025)
Dynamic reconfiguration for malleable applications using RMA
di: Martín-Álvarez, Iker, et al.
Pubblicazione: (2025)
di: Martín-Álvarez, Iker, et al.
Pubblicazione: (2025)
A reliability- and latency-driven task allocation framework for workflow applications in the edge-hub-cloud continuum
di: Kouloumpris, Andreas, et al.
Pubblicazione: (2026)
di: Kouloumpris, Andreas, et al.
Pubblicazione: (2026)
Perpetual Exploration of a Ring in Presence of Byzantine Black Hole
di: Goswami, Pritam, et al.
Pubblicazione: (2024)
di: Goswami, Pritam, et al.
Pubblicazione: (2024)
Run-time application migration using checkpoint/restore in userspace
di: Tošić, Aleksandar
Pubblicazione: (2023)
di: Tošić, Aleksandar
Pubblicazione: (2023)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
di: Jiang, Xuan, et al.
Pubblicazione: (2024)
di: Jiang, Xuan, et al.
Pubblicazione: (2024)
Research on fault diagnosis and root cause analysis based on full stack observability
di: Hou, Jian
Pubblicazione: (2025)
di: Hou, Jian
Pubblicazione: (2025)
From Patchwork to Network: A Comprehensive Framework for Demand Analysis and Fleet Optimization of Urban Air Mobility
di: Jiang, Xuan, et al.
Pubblicazione: (2025)
di: Jiang, Xuan, et al.
Pubblicazione: (2025)
Multi-objective application placement in fog computing using graph neural network-based reinforcement learning
di: Lera, Isaac, et al.
Pubblicazione: (2026)
di: Lera, Isaac, et al.
Pubblicazione: (2026)
Towards observability of scientific applications
di: Balis, Bartosz, et al.
Pubblicazione: (2024)
di: Balis, Bartosz, et al.
Pubblicazione: (2024)
Configuration management in the distributed cloud
di: Ranković, Tamara, et al.
Pubblicazione: (2024)
di: Ranković, Tamara, et al.
Pubblicazione: (2024)
FailSafe: High-performance Resilient Serving
di: Xu, Ziyi, et al.
Pubblicazione: (2025)
di: Xu, Ziyi, et al.
Pubblicazione: (2025)
Software engineering to sustain a high-performance computing scientific application: QMCPACK
di: Godoy, William F., et al.
Pubblicazione: (2023)
di: Godoy, William F., et al.
Pubblicazione: (2023)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
A platform for lightweight deployment of IoT applications based on a Function-as-a-Service model
di: Sansó, Sebastià, et al.
Pubblicazione: (2024)
di: Sansó, Sebastià, et al.
Pubblicazione: (2024)
Accelerating discovery across scientific disciplines through reproducible workflows with AiiDAlab
di: Yakutovich, Aliaksandr V., et al.
Pubblicazione: (2025)
di: Yakutovich, Aliaksandr V., et al.
Pubblicazione: (2025)
Hierarchical storage management in user space for neuroimaging applications
di: Hayot-Sasson, Valérie, et al.
Pubblicazione: (2024)
di: Hayot-Sasson, Valérie, et al.
Pubblicazione: (2024)
On the performance of two-sided MPI, MPI-3 RMA and SHMEM in a Lagrangian particle cluster algorithm
di: Frey, Matthias, et al.
Pubblicazione: (2024)
di: Frey, Matthias, et al.
Pubblicazione: (2024)
Boosting performance: Gradient Clock Synchronisation with two-way measured links
di: Wenning, Sophie
Pubblicazione: (2025)
di: Wenning, Sophie
Pubblicazione: (2025)
Scrutiny new framework in integrated distributed reliable systems
di: Gashti, Mehdi Zekriyapanah
Pubblicazione: (2025)
di: Gashti, Mehdi Zekriyapanah
Pubblicazione: (2025)
CRIU -- Checkpoint Restore in Userspace for computational simulations and scientific applications
di: Andrijauskas, Fabio, et al.
Pubblicazione: (2024)
di: Andrijauskas, Fabio, et al.
Pubblicazione: (2024)
Proxima. A DAG based cooperative distributed ledger
di: Drasutis, Evaldas
Pubblicazione: (2024)
di: Drasutis, Evaldas
Pubblicazione: (2024)
Fog enabled distributed training architecture for federated learning
di: Kumar, Aditya, et al.
Pubblicazione: (2024)
di: Kumar, Aditya, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Visualizing Distributed Traces in Aggregate
di: Samanta, Adrita, et al.
Pubblicazione: (2024) -
FIRED: a fine-grained robust performance diagnosis framework for cloud applications
di: Xin, Ruyue, et al.
Pubblicazione: (2022) -
Cloud abstractions for AI workloads
di: Canini, Marco, et al.
Pubblicazione: (2025) -
emucxl: an emulation framework for CXL-based disaggregated memory applications
di: Gond, Raja, et al.
Pubblicazione: (2024) -
Portable, heterogeneous ensemble workflows at scale using libEnsemble
di: Hudson, Stephen, et al.
Pubblicazione: (2024)