STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chen, Liu, Liwei, Tao, Jun, Yang, Xiaoyu, Xu, Xuenan, Chen, Kai, Zhou, Bowen, Wu, Wen, Zhang, Chao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914408877260800
author Zhang, Chen
Liu, Liwei
Tao, Jun
Yang, Xiaoyu
Xu, Xuenan
Chen, Kai
Zhou, Bowen
Wu, Wen
Zhang, Chao
author_facet Zhang, Chen
Liu, Liwei
Tao, Jun
Yang, Xiaoyu
Xu, Xuenan
Chen, Kai
Zhou, Bowen
Wu, Wen
Zhang, Chao
contents Scientific time series are central to scientific AI but are typically sparse, highly heterogeneous, and limited in scale, making unified representation learning particularly challenging. Meanwhile, foundation models pretrained on relevant time series domains such as audio, general time series, and brain signals contain rich knowledge, but their applicability to scientific signals remains underexplored. In this paper, we investigate the transferability and complementarity of foundation models from relevant time series domains, and study how to effectively leverage them to build a unified encoder for scientific time series. We first systematically evaluate relevant foundation models, showing the effectiveness of knowledge transfer to scientific tasks and their complementary strengths. Based on this observation, we propose STEP, a Scientific Time Series Encoder Pretraining framework via cross domain distillation. STEP introduces adaptive patching to handle extreme-length sequences and a statistics compensation scheme to accommodate diverse numerical scales. It further leverages cross-domain distillation to integrate knowledge from multiple foundation models into a unified encoder. By combining complementary representations across different domains, STEP learns general-purpose and transferable features tailored for scientific signals. Experiments on seven scientific time series tasks demonstrate that STEP provides both an effective structure and an effective pretraining paradigm, taking a STEP toward scientific time series representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18688
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation
Zhang, Chen
Liu, Liwei
Tao, Jun
Yang, Xiaoyu
Xu, Xuenan
Chen, Kai
Zhou, Bowen
Wu, Wen
Zhang, Chao
Machine Learning
Computation and Language
Scientific time series are central to scientific AI but are typically sparse, highly heterogeneous, and limited in scale, making unified representation learning particularly challenging. Meanwhile, foundation models pretrained on relevant time series domains such as audio, general time series, and brain signals contain rich knowledge, but their applicability to scientific signals remains underexplored. In this paper, we investigate the transferability and complementarity of foundation models from relevant time series domains, and study how to effectively leverage them to build a unified encoder for scientific time series. We first systematically evaluate relevant foundation models, showing the effectiveness of knowledge transfer to scientific tasks and their complementary strengths. Based on this observation, we propose STEP, a Scientific Time Series Encoder Pretraining framework via cross domain distillation. STEP introduces adaptive patching to handle extreme-length sequences and a statistics compensation scheme to accommodate diverse numerical scales. It further leverages cross-domain distillation to integrate knowledge from multiple foundation models into a unified encoder. By combining complementary representations across different domains, STEP learns general-purpose and transferable features tailored for scientific signals. Experiments on seven scientific time series tasks demonstrate that STEP provides both an effective structure and an effective pretraining paradigm, taking a STEP toward scientific time series representation learning.
title STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2603.18688