FIRED: a fine-grained robust performance diagnosis framework for cloud applications
Fuente:
arXiv
Guardado en:
| Autores principales: | Xin, Ruyue, Liu, Hongyun, Chen, Peng, Grosso, Paola, Zhao, Zhiming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Privacy-, Budget-, and Deadline-Aware Service Optimization for Large Medical Image Processing across Hybrid Clouds
por: Wang, Yuandou, et al.
Publicado: (2024)
por: Wang, Yuandou, et al.
Publicado: (2024)
The workflow motif: a widely-useful performance diagnosis abstraction for distributed applications
por: Abdi, Mania, et al.
Publicado: (2025)
por: Abdi, Mania, et al.
Publicado: (2025)
Dflow, a Python framework for constructing cloud-native AI-for-Science workflows
por: Liu, Xinzijian, et al.
Publicado: (2024)
por: Liu, Xinzijian, et al.
Publicado: (2024)
Managing Federated Learning on Decentralized Infrastructures as a Reputation-based Collaborative Workflow
por: Wang, Yuandou, et al.
Publicado: (2025)
por: Wang, Yuandou, et al.
Publicado: (2025)
Misconfiguration prevention and error cause detection for distributed-cloud applications
por: Ranković, Tamara, et al.
Publicado: (2024)
por: Ranković, Tamara, et al.
Publicado: (2024)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
por: Wu, Jingfeng, et al.
Publicado: (2024)
por: Wu, Jingfeng, et al.
Publicado: (2024)
D-VRE: From a Jupyter-enabled Private Research Environment to Decentralized Collaborative Research Ecosystem
por: Wang, Yuandou, et al.
Publicado: (2024)
por: Wang, Yuandou, et al.
Publicado: (2024)
A study of the spectrum resource leasing method based on ERC4907 extension
por: Liang, Zhiming, et al.
Publicado: (2025)
por: Liang, Zhiming, et al.
Publicado: (2025)
emucxl: an emulation framework for CXL-based disaggregated memory applications
por: Gond, Raja, et al.
Publicado: (2024)
por: Gond, Raja, et al.
Publicado: (2024)
A reliability- and latency-driven task allocation framework for workflow applications in the edge-hub-cloud continuum
por: Kouloumpris, Andreas, et al.
Publicado: (2026)
por: Kouloumpris, Andreas, et al.
Publicado: (2026)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
por: Zhao, Yihao, et al.
Publicado: (2025)
por: Zhao, Yihao, et al.
Publicado: (2025)
Configuration management in the distributed cloud
por: Ranković, Tamara, et al.
Publicado: (2024)
por: Ranković, Tamara, et al.
Publicado: (2024)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
por: Li, Cong, et al.
Publicado: (2025)
por: Li, Cong, et al.
Publicado: (2025)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
por: Schieffer, Gabin, et al.
Publicado: (2026)
por: Schieffer, Gabin, et al.
Publicado: (2026)
HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs
por: Li, Yanliang, et al.
Publicado: (2025)
por: Li, Yanliang, et al.
Publicado: (2025)
Towards Seamless Serverless Computing Across an Edge-Cloud Continuum
por: Simion, Emilian, et al.
Publicado: (2024)
por: Simion, Emilian, et al.
Publicado: (2024)
Trace-based, time-resolved analysis of MPI application performance using standard metrics
por: Haldar, Kingshuk
Publicado: (2025)
por: Haldar, Kingshuk
Publicado: (2025)
Fine-grained MoE Load Balancing with Linear Programming
por: Zhao, Chenqi, et al.
Publicado: (2025)
por: Zhao, Chenqi, et al.
Publicado: (2025)
Towards cloud-native scientific workflow management
por: Orzechowski, Michal, et al.
Publicado: (2024)
por: Orzechowski, Michal, et al.
Publicado: (2024)
TALP-Pages: An easy-to-integrate continuous performance monitoring framework
por: Seitz, Valentin, et al.
Publicado: (2025)
por: Seitz, Valentin, et al.
Publicado: (2025)
A monitoring system for collecting and aggregating metrics from distributed clouds
por: Ranković, Tamara, et al.
Publicado: (2026)
por: Ranković, Tamara, et al.
Publicado: (2026)
Enabling an OpenStack-based cloud on top of RISC-V hardware
por: Marrón, Diego, et al.
Publicado: (2024)
por: Marrón, Diego, et al.
Publicado: (2024)
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
por: Ni, Xiaobing, et al.
Publicado: (2024)
por: Ni, Xiaobing, et al.
Publicado: (2024)
Multi-agent Reinforcement Learning-based In-place Scaling Engine for Edge-cloud Systems
por: Prodanov, Jovan, et al.
Publicado: (2025)
por: Prodanov, Jovan, et al.
Publicado: (2025)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
por: Shi, Ruimin, et al.
Publicado: (2026)
por: Shi, Ruimin, et al.
Publicado: (2026)
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
por: Mei, Junyi, et al.
Publicado: (2024)
por: Mei, Junyi, et al.
Publicado: (2024)
Adaptive multi-criteria-based load balancing technique for resource allocation in fog-cloud environments
por: Gad-Elrab, Ahmed A. A., et al.
Publicado: (2024)
por: Gad-Elrab, Ahmed A. A., et al.
Publicado: (2024)
Research on fault diagnosis and root cause analysis based on full stack observability
por: Hou, Jian
Publicado: (2025)
por: Hou, Jian
Publicado: (2025)
Exploring Fine-grained Task Parallelism on Simultaneous Multithreading Cores
por: Los, Denis, et al.
Publicado: (2024)
por: Los, Denis, et al.
Publicado: (2024)
An optimization framework for task allocation in the edge/hub/cloud paradigm
por: Kouloumpris, Andreas, et al.
Publicado: (2025)
por: Kouloumpris, Andreas, et al.
Publicado: (2025)
From SLA to vendor-neutral metrics: An intelligent knowledge-based approach for multi-cloud SLA-based broker
por: Rampérez, Víctor, et al.
Publicado: (2025)
por: Rampérez, Víctor, et al.
Publicado: (2025)
Performance analysis of mdx II: A next-generation cloud platform for cross-disciplinary data science research
por: Takahashi, Keichi, et al.
Publicado: (2025)
por: Takahashi, Keichi, et al.
Publicado: (2025)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
por: Tsolakis, Christos, et al.
Publicado: (2024)
por: Tsolakis, Christos, et al.
Publicado: (2024)
A common parallel framework for LLP combinatorial problems
por: Alves, David Ribeiro, et al.
Publicado: (2026)
por: Alves, David Ribeiro, et al.
Publicado: (2026)
Scrutiny new framework in integrated distributed reliable systems
por: Gashti, Mehdi Zekriyapanah
Publicado: (2025)
por: Gashti, Mehdi Zekriyapanah
Publicado: (2025)
Data-Locality-Aware Task Assignment and Scheduling for Distributed Job Executions
por: Zhao, Hailiang, et al.
Publicado: (2024)
por: Zhao, Hailiang, et al.
Publicado: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
por: Mo, Zizhao, et al.
Publicado: (2025)
por: Mo, Zizhao, et al.
Publicado: (2025)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
por: Wu, Jingfeng, et al.
Publicado: (2025)
por: Wu, Jingfeng, et al.
Publicado: (2025)
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
por: Hu, Bodun, et al.
Publicado: (2024)
por: Hu, Bodun, et al.
Publicado: (2024)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
por: Gu, Diandian, et al.
Publicado: (2024)
por: Gu, Diandian, et al.
Publicado: (2024)
Ejemplares similares
-
Towards Privacy-, Budget-, and Deadline-Aware Service Optimization for Large Medical Image Processing across Hybrid Clouds
por: Wang, Yuandou, et al.
Publicado: (2024) -
The workflow motif: a widely-useful performance diagnosis abstraction for distributed applications
por: Abdi, Mania, et al.
Publicado: (2025) -
Dflow, a Python framework for constructing cloud-native AI-for-Science workflows
por: Liu, Xinzijian, et al.
Publicado: (2024) -
Managing Federated Learning on Decentralized Infrastructures as a Reputation-based Collaborative Workflow
por: Wang, Yuandou, et al.
Publicado: (2025) -
Misconfiguration prevention and error cause detection for distributed-cloud applications
por: Ranković, Tamara, et al.
Publicado: (2024)