TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910256943071232 |
|---|---|
| author | Li, Hongkai Xie, Shifeng Shen, Lefei Li, Zhuo Chen, Mouxiang Zhang, Xiaobin Fu, Han Sun, Jianling Ren, Xiaoxue Liu, Chenghao |
| author_facet | Li, Hongkai Xie, Shifeng Shen, Lefei Li, Zhuo Chen, Mouxiang Zhang, Xiaobin Fu, Han Sun, Jianling Ren, Xiaoxue Liu, Chenghao |
| contents | Time series foundation models (TSFMs) are increasingly pretrained on large corpora, raising concerns that evaluation datasets may have been exposed during pretraining and thus yield overly optimistic performance estimates. Auditing such contamination is challenging in time series because signals are continuous and heterogeneous, and often lack corpus documentation. To the best of our knowledge, this is the first work to study pretraining contamination auditing for TSFMs. We formalize the problem of pretraining contamination auditing for TSFMs and propose TSFMAudit, a method based on probe adaptation dynamics. Our key intuition is that contamination manifests as unusually efficient adaptation: after a fine tuning probe, contaminated datasets tend to exhibit faster loss reduction with smaller backbone movement. We evaluate TSFMAudit on 6 TSFMs and 187 datasets using documented training source evidence as supervision, and compare against 10 competitive baselines adapted from the LLM literature. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_26161 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models Li, Hongkai Xie, Shifeng Shen, Lefei Li, Zhuo Chen, Mouxiang Zhang, Xiaobin Fu, Han Sun, Jianling Ren, Xiaoxue Liu, Chenghao Machine Learning Artificial Intelligence Time series foundation models (TSFMs) are increasingly pretrained on large corpora, raising concerns that evaluation datasets may have been exposed during pretraining and thus yield overly optimistic performance estimates. Auditing such contamination is challenging in time series because signals are continuous and heterogeneous, and often lack corpus documentation. To the best of our knowledge, this is the first work to study pretraining contamination auditing for TSFMs. We formalize the problem of pretraining contamination auditing for TSFMs and propose TSFMAudit, a method based on probe adaptation dynamics. Our key intuition is that contamination manifests as unusually efficient adaptation: after a fine tuning probe, contaminated datasets tend to exhibit faster loss reduction with smaller backbone movement. We evaluate TSFMAudit on 6 TSFMs and 187 datasets using documented training source evidence as supervision, and compare against 10 competitive baselines adapted from the LLM literature. |
| title | TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2605.26161 |