When to Retrain after Drift: A Data-Only Test of Post-Drift Data Size Sufficiency

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fujiwara, Ren, Matsubara, Yasuko, Sakurai, Yasushi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914582994354176
author Fujiwara, Ren
Matsubara, Yasuko
Sakurai, Yasushi
author_facet Fujiwara, Ren
Matsubara, Yasuko
Sakurai, Yasushi
contents Sudden concept drift makes previously trained predictors unreliable, yet deciding when to retrain and what post-drift data size is sufficient is rarely addressed. We propose CALIPER - a detector- and model-agnostic, data-only test that estimates the post-drift data size required for stable retraining. CALIPER exploits state dependence in streams generated by dynamical systems: we run a single-pass weighted local regression over the post-drift window and track a one-step proxy error as a function of a locality parameter $θ$. When an effective sample size gate is satisfied, a monotonically non-increasing trend in this error with increasing a locality parameter indicates that the data size is sufficiently informative for retraining. We also provide a theoretical analysis of our method, and we show that the algorithm has a low per-update time and memory. Across datasets from four heterogeneous domains, three learner families, and two detectors, CALIPER consistently matches or exceeds the best fixed data size for retraining while incurring negligible overhead and often outperforming incremental updates. CALIPER closes the gap between drift detection and data-sufficient adaptation in streaming learning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09024
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When to Retrain after Drift: A Data-Only Test of Post-Drift Data Size Sufficiency
Fujiwara, Ren
Matsubara, Yasuko
Sakurai, Yasushi
Machine Learning
Sudden concept drift makes previously trained predictors unreliable, yet deciding when to retrain and what post-drift data size is sufficient is rarely addressed. We propose CALIPER - a detector- and model-agnostic, data-only test that estimates the post-drift data size required for stable retraining. CALIPER exploits state dependence in streams generated by dynamical systems: we run a single-pass weighted local regression over the post-drift window and track a one-step proxy error as a function of a locality parameter $θ$. When an effective sample size gate is satisfied, a monotonically non-increasing trend in this error with increasing a locality parameter indicates that the data size is sufficiently informative for retraining. We also provide a theoretical analysis of our method, and we show that the algorithm has a low per-update time and memory. Across datasets from four heterogeneous domains, three learner families, and two detectors, CALIPER consistently matches or exceeds the best fixed data size for retraining while incurring negligible overhead and often outperforming incremental updates. CALIPER closes the gap between drift detection and data-sufficient adaptation in streaming learning.
title When to Retrain after Drift: A Data-Only Test of Post-Drift Data Size Sufficiency
topic Machine Learning
url https://arxiv.org/abs/2603.09024