VITRO: Vocabulary Inversion for Time-series Representation Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bellos, Filippos, Nguyen, Nam H., Corso, Jason J.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909439918866432
author Bellos, Filippos
Nguyen, Nam H.
Corso, Jason J.
author_facet Bellos, Filippos
Nguyen, Nam H.
Corso, Jason J.
contents Although LLMs have demonstrated remarkable capabilities in processing and generating textual data, their pre-trained vocabularies are ill-suited for capturing the nuanced temporal dynamics and patterns inherent in time series. The discrete, symbolic nature of natural language tokens, which these vocabularies are designed to represent, does not align well with the continuous, numerical nature of time series data. To address this fundamental limitation, we propose VITRO. Our method adapts textual inversion optimization from the vision-language domain in order to learn a new time series per-dataset vocabulary that bridges the gap between the discrete, semantic nature of natural language and the continuous, numerical nature of time series data. We show that learnable time series-specific pseudo-word embeddings represent time series data better than existing general language model vocabularies, with VITRO-enhanced methods achieving state-of-the-art performance in long-term forecasting across most datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17921
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VITRO: Vocabulary Inversion for Time-series Representation Optimization
Bellos, Filippos
Nguyen, Nam H.
Corso, Jason J.
Machine Learning
Computation and Language
Although LLMs have demonstrated remarkable capabilities in processing and generating textual data, their pre-trained vocabularies are ill-suited for capturing the nuanced temporal dynamics and patterns inherent in time series. The discrete, symbolic nature of natural language tokens, which these vocabularies are designed to represent, does not align well with the continuous, numerical nature of time series data. To address this fundamental limitation, we propose VITRO. Our method adapts textual inversion optimization from the vision-language domain in order to learn a new time series per-dataset vocabulary that bridges the gap between the discrete, semantic nature of natural language and the continuous, numerical nature of time series data. We show that learnable time series-specific pseudo-word embeddings represent time series data better than existing general language model vocabularies, with VITRO-enhanced methods achieving state-of-the-art performance in long-term forecasting across most datasets.
title VITRO: Vocabulary Inversion for Time-series Representation Optimization
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2412.17921