Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Fang, Huang, Yuyang, Yu, Miao, Ma, Sixiang, Liu, Tongping, Wang, Yang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917205676916736
author Zhou, Fang
Huang, Yuyang
Yu, Miao
Ma, Sixiang
Liu, Tongping
Wang, Yang
author_facet Zhou, Fang
Huang, Yuyang
Yu, Miao
Ma, Sixiang
Liu, Tongping
Wang, Yang
contents This paper presents our experience to understand latency variance caused by kernel and hardware events, which are often invisible at the application level. For this purpose, we have built VarMRI, a tool chain to monitor and analyze those events in the long term. To mitigate the "big data" problem caused by long-term monitoring, VarMRI selectively records a subset of events following two principles: it only records events that are affecting the requests recorded by the application; it records coarse-grained information first and records additional information only when necessary. Furthermore, VarMRI introduces an analysis method that is efficient on large amount of data, robust on different data set and against missing data, and informative to the user. VarMRI has helped us to carry out a 3,000-hour study of six applications and benchmarks on CloudLab. It reveals a wide variety of events causing latency variance, including interrupt preemption, Java GC, pipeline stall, NUMA balancing etc.; simple optimization or tuning can reduce tail latencies by up to 31%. Furthermore, the impacts of some of these events vary significantly across different experiments, which confirms the necessity of long-term monitoring.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10572
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
Zhou, Fang
Huang, Yuyang
Yu, Miao
Ma, Sixiang
Liu, Tongping
Wang, Yang
Performance
This paper presents our experience to understand latency variance caused by kernel and hardware events, which are often invisible at the application level. For this purpose, we have built VarMRI, a tool chain to monitor and analyze those events in the long term. To mitigate the "big data" problem caused by long-term monitoring, VarMRI selectively records a subset of events following two principles: it only records events that are affecting the requests recorded by the application; it records coarse-grained information first and records additional information only when necessary. Furthermore, VarMRI introduces an analysis method that is efficient on large amount of data, robust on different data set and against missing data, and informative to the user. VarMRI has helped us to carry out a 3,000-hour study of six applications and benchmarks on CloudLab. It reveals a wide variety of events causing latency variance, including interrupt preemption, Java GC, pipeline stall, NUMA balancing etc.; simple optimization or tuning can reduce tail latencies by up to 31%. Furthermore, the impacts of some of these events vary significantly across different experiments, which confirms the necessity of long-term monitoring.
title Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
topic Performance
url https://arxiv.org/abs/2601.10572