Byte Pair Encoding for Efficient Time Series Forecasting

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Götz, Leon, Kollovieh, Marcel, Günnemann, Stephan, Schwinn, Leo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911740044771328
author Götz, Leon
Kollovieh, Marcel
Günnemann, Stephan
Schwinn, Leo
author_facet Götz, Leon
Kollovieh, Marcel
Günnemann, Stephan
Schwinn, Leo
contents Existing time series tokenization methods predominantly encode a constant number of samples into individual tokens. This inflexible approach can generate excessive tokens for even simple patterns like extended constant values, resulting in substantial computational overhead. Inspired by the success of byte pair encoding, we propose the first pattern-centric tokenization scheme for time series analysis. Based on a discrete vocabulary of frequent motifs, our method merges samples with underlying patterns into tokens, compressing time series adaptively. Exploiting our finite set of motifs and the continuous properties of time series, we further introduce conditional decoding as a lightweight yet powerful post-hoc optimization method, which requires no gradient computation and adds no computational overhead. On recent time series foundation models, our motif-based tokenization improves forecasting performance by 40% and boosts efficiency by 2314% on average. Conditional decoding further reduces MSE by up to 48%. In an extensive analysis, we demonstrate the adaptiveness of our tokenization to diverse temporal patterns, its generalization to unseen data, and its meaningful token representations capturing distinct time series properties, including statistical moments and trends.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14411
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Byte Pair Encoding for Efficient Time Series Forecasting
Götz, Leon
Kollovieh, Marcel
Günnemann, Stephan
Schwinn, Leo
Machine Learning
Existing time series tokenization methods predominantly encode a constant number of samples into individual tokens. This inflexible approach can generate excessive tokens for even simple patterns like extended constant values, resulting in substantial computational overhead. Inspired by the success of byte pair encoding, we propose the first pattern-centric tokenization scheme for time series analysis. Based on a discrete vocabulary of frequent motifs, our method merges samples with underlying patterns into tokens, compressing time series adaptively. Exploiting our finite set of motifs and the continuous properties of time series, we further introduce conditional decoding as a lightweight yet powerful post-hoc optimization method, which requires no gradient computation and adds no computational overhead. On recent time series foundation models, our motif-based tokenization improves forecasting performance by 40% and boosts efficiency by 2314% on average. Conditional decoding further reduces MSE by up to 48%. In an extensive analysis, we demonstrate the adaptiveness of our tokenization to diverse temporal patterns, its generalization to unseen data, and its meaningful token representations capturing distinct time series properties, including statistical moments and trends.
title Byte Pair Encoding for Efficient Time Series Forecasting
topic Machine Learning
url https://arxiv.org/abs/2505.14411