PANDA: Noise-Resilient Antagonist Identification in Production Datacenters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Sixiang, Deng, Nan, Rzadca, Krzysiek, Lin, Xiaojun, Hu, Y. Charlie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914168041373696
author Zhou, Sixiang
Deng, Nan
Rzadca, Krzysiek
Lin, Xiaojun
Hu, Y. Charlie
author_facet Zhou, Sixiang
Deng, Nan
Rzadca, Krzysiek
Lin, Xiaojun
Hu, Y. Charlie
contents Modern warehouse-scale datacenters commonly collocate multiple jobs on shared machines to improve resource utilization. However, such collocation often leads to performance interference caused by antagonistic jobs that overconsume shared resources. Existing antagonist-detection approaches either rely on offline profiling, which is costly and unscalable, or use a sample-from-production approach, which suffers from noisy measurements and fails under multi-victim scenarios. We present PANDA, a noise-resilient antagonist identification framework for production-scale datacenters. Like prior correlation-based methods, PANDA uses cycles per instruction (CPI) as its performance metric, but it differs by (i) leveraging global historical knowledge across all machines to suppress sampling noise and (ii) introducing a machine-level CPI metric that captures shared-resource contention among multiple co-located tasks. Evaluation on a recent Google production trace shows that PANDA ranks true antagonists far more accurately than prior methods -- improving average suspicion percentile from 50-55% to 82.6% -- and achieves consistent antagonist identification under multi-victim scenarios, all with negligible runtime overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PANDA: Noise-Resilient Antagonist Identification in Production Datacenters
Zhou, Sixiang
Deng, Nan
Rzadca, Krzysiek
Lin, Xiaojun
Hu, Y. Charlie
Performance
Modern warehouse-scale datacenters commonly collocate multiple jobs on shared machines to improve resource utilization. However, such collocation often leads to performance interference caused by antagonistic jobs that overconsume shared resources. Existing antagonist-detection approaches either rely on offline profiling, which is costly and unscalable, or use a sample-from-production approach, which suffers from noisy measurements and fails under multi-victim scenarios. We present PANDA, a noise-resilient antagonist identification framework for production-scale datacenters. Like prior correlation-based methods, PANDA uses cycles per instruction (CPI) as its performance metric, but it differs by (i) leveraging global historical knowledge across all machines to suppress sampling noise and (ii) introducing a machine-level CPI metric that captures shared-resource contention among multiple co-located tasks. Evaluation on a recent Google production trace shows that PANDA ranks true antagonists far more accurately than prior methods -- improving average suspicion percentile from 50-55% to 82.6% -- and achieves consistent antagonist identification under multi-victim scenarios, all with negligible runtime overhead.
title PANDA: Noise-Resilient Antagonist Identification in Production Datacenters
topic Performance
url https://arxiv.org/abs/2511.08803