Reliable Microservice Tail Latency Prediction via Decoupled Dual-Stream Learning and Gradient Modulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Qian, Wenzhuo, Zhao, Hailiang, Chen, Jiayi, Wang, Ziqi, Chen, Tianlv, Ling, Zhiwei, Zhao, Xinkui, Chow, Kingsum, Zomaya, Albert Y., Deng, Shuiguang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveraging
di: Ling, Zhiwei, et al.
Pubblicazione: (2025)
di: Ling, Zhiwei, et al.
Pubblicazione: (2025)
Scene-Aware Latency Estimation for Microservices via Multi-Scale Graph Fusion
di: Sun, Zhichao, et al.
Pubblicazione: (2026)
di: Sun, Zhichao, et al.
Pubblicazione: (2026)
Missing-Aware Multimodal Fusion for Unified Microservice Incident Management
di: Qian, Wenzhuo, et al.
Pubblicazione: (2026)
di: Qian, Wenzhuo, et al.
Pubblicazione: (2026)
Morphis: SLO-Aware Resource Scheduling for Microservices with Time-Varying Call Graphs
di: Tang, Yu, et al.
Pubblicazione: (2026)
di: Tang, Yu, et al.
Pubblicazione: (2026)
Agentic Services Computing
di: Deng, Shuiguang, et al.
Pubblicazione: (2025)
di: Deng, Shuiguang, et al.
Pubblicazione: (2025)
Adaptive Dual-Weighting Framework for Federated Learning via Out-of-Distribution Detection
di: Ling, Zhiwei, et al.
Pubblicazione: (2026)
di: Ling, Zhiwei, et al.
Pubblicazione: (2026)
Employing Software Diversity in Cloud Microservices to Engineer Reliable and Performant Systems
di: Akhtarian, Nazanin, et al.
Pubblicazione: (2024)
di: Akhtarian, Nazanin, et al.
Pubblicazione: (2024)
An Experimental Study of Low-Latency Video Streaming over 5G
di: Khan, Imran, et al.
Pubblicazione: (2024)
di: Khan, Imran, et al.
Pubblicazione: (2024)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
di: Jeffery, Andrew, et al.
Pubblicazione: (2024)
di: Jeffery, Andrew, et al.
Pubblicazione: (2024)
Root Cause Localization for Microservice Systems in Cloud-edge Collaborative Environments
di: Zhu, Yuhan, et al.
Pubblicazione: (2024)
di: Zhu, Yuhan, et al.
Pubblicazione: (2024)
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
di: Sun, Jiaqi, et al.
Pubblicazione: (2025)
di: Sun, Jiaqi, et al.
Pubblicazione: (2025)
Strongly Tail-Optimal Scheduling in the Light-Tailed M/G/1
di: Yu, George, et al.
Pubblicazione: (2024)
di: Yu, George, et al.
Pubblicazione: (2024)
When Does Hierarchy Help? Benchmarking Agent Coordination in Event-Driven Industrial Scheduling
di: Wang, Ziqi, et al.
Pubblicazione: (2026)
di: Wang, Ziqi, et al.
Pubblicazione: (2026)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
di: Shen, Siyuan, et al.
Pubblicazione: (2025)
di: Shen, Siyuan, et al.
Pubblicazione: (2025)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
di: Ahmadvand, Hossein, et al.
Pubblicazione: (2025)
di: Ahmadvand, Hossein, et al.
Pubblicazione: (2025)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
di: Zhou, Fang, et al.
Pubblicazione: (2026)
di: Zhou, Fang, et al.
Pubblicazione: (2026)
A Model-driven Approach for Continuous Performance Engineering in Microservice-based Systems
di: Cortellessa, Vittorio, et al.
Pubblicazione: (2023)
di: Cortellessa, Vittorio, et al.
Pubblicazione: (2023)
An Empirical Study on How Architectural Topology Affects Microservice Performance and Energy Usage
di: Ristova, Irena, et al.
Pubblicazione: (2026)
di: Ristova, Irena, et al.
Pubblicazione: (2026)
A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
di: Sapkas, M., et al.
Pubblicazione: (2025)
di: Sapkas, M., et al.
Pubblicazione: (2025)
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
di: Le, Truong-Thanh, et al.
Pubblicazione: (2026)
di: Le, Truong-Thanh, et al.
Pubblicazione: (2026)
Motion-to-Motion Latency Measurement Framework for Connected and Autonomous Vehicle Teleoperation
di: Provost, François, et al.
Pubblicazione: (2025)
di: Provost, François, et al.
Pubblicazione: (2025)
SPLIT: SymPathy for Large jobs Improves Tail latency
di: Li, Zhouzi, et al.
Pubblicazione: (2026)
di: Li, Zhouzi, et al.
Pubblicazione: (2026)
Tail Optimality and Performance Analysis of the Nudge*(M) Scheduling Algorithm
di: Charlet, Nils, et al.
Pubblicazione: (2024)
di: Charlet, Nils, et al.
Pubblicazione: (2024)
Streaming Data in HPC Workflows Using ADIOS
di: Eisenhauer, Greg, et al.
Pubblicazione: (2024)
di: Eisenhauer, Greg, et al.
Pubblicazione: (2024)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
di: Chen, David, et al.
Pubblicazione: (2026)
di: Chen, David, et al.
Pubblicazione: (2026)
Latency Based Tiling
di: Cashman, Jack
Pubblicazione: (2025)
di: Cashman, Jack
Pubblicazione: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
di: Kong, Linghao, et al.
Pubblicazione: (2026)
di: Kong, Linghao, et al.
Pubblicazione: (2026)
TraDE: Network and Traffic-aware Adaptive Scheduling for Microservices Under Dynamics
di: Chen, Ming, et al.
Pubblicazione: (2024)
di: Chen, Ming, et al.
Pubblicazione: (2024)
Towards Multi-dimensional Elasticity for Pervasive Stream Processing Services
di: Sedlak, Boris, et al.
Pubblicazione: (2025)
di: Sedlak, Boris, et al.
Pubblicazione: (2025)
Tail Bounds for Queues with Abandonment: Constant, Moderate, Large Deviations, and Efficient Concentration
di: Wang, Zedong, et al.
Pubblicazione: (2026)
di: Wang, Zedong, et al.
Pubblicazione: (2026)
When Does the Gittins Policy Have Asymptotically Optimal Response Time Tail?
di: Scully, Ziv, et al.
Pubblicazione: (2021)
di: Scully, Ziv, et al.
Pubblicazione: (2021)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
di: Zhao, Yushang, et al.
Pubblicazione: (2025)
di: Zhao, Yushang, et al.
Pubblicazione: (2025)
Optimizing Stateful Microservice Migration in Kubernetes with MS2M and Forensic Checkpointing
di: Dinh-Tuan, Hai, et al.
Pubblicazione: (2025)
di: Dinh-Tuan, Hai, et al.
Pubblicazione: (2025)
Towards A Flexible Accuracy-Oriented Deep Learning Module Inference Latency Prediction Framework for Adaptive Optimization Algorithms
di: Shen, Jingran, et al.
Pubblicazione: (2023)
di: Shen, Jingran, et al.
Pubblicazione: (2023)
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
di: Bando, Yosuke, et al.
Pubblicazione: (2025)
di: Bando, Yosuke, et al.
Pubblicazione: (2025)
Do LLMs Have Visualization Literacy? An Evaluation on Modified Visualizations to Test Generalization in Data Interpretation
di: Hong, Jiayi, et al.
Pubblicazione: (2025)
di: Hong, Jiayi, et al.
Pubblicazione: (2025)
Colored Markov Modulated Fluid Queues
di: Van Houdt, Benny
Pubblicazione: (2026)
di: Van Houdt, Benny
Pubblicazione: (2026)
Revealing NVIDIA Closed-Source Driver Command Streams for CPU-GPU Runtime Behavior Insight
di: Yan, Yuang, et al.
Pubblicazione: (2026)
di: Yan, Yuang, et al.
Pubblicazione: (2026)
Energy Efficiency Analysis of Active RIS-enhanced Wireless Network under Power-Sum Constraint
di: Xin, Jingdie, et al.
Pubblicazione: (2025)
di: Xin, Jingdie, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveraging
di: Ling, Zhiwei, et al.
Pubblicazione: (2025) -
Scene-Aware Latency Estimation for Microservices via Multi-Scale Graph Fusion
di: Sun, Zhichao, et al.
Pubblicazione: (2026) -
Missing-Aware Multimodal Fusion for Unified Microservice Incident Management
di: Qian, Wenzhuo, et al.
Pubblicazione: (2026) -
Morphis: SLO-Aware Resource Scheduling for Microservices with Time-Varying Call Graphs
di: Tang, Yu, et al.
Pubblicazione: (2026) -
Agentic Services Computing
di: Deng, Shuiguang, et al.
Pubblicazione: (2025)