Chronicle: A Multimodal Foundation Model for Joint Language and Time Series Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Quinlan, Paul, Levasseur, Jeremy, Li, Qingguo, Zhu, Xiaodan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911698826297344
author Quinlan, Paul
Levasseur, Jeremy
Li, Qingguo
Zhu, Xiaodan
author_facet Quinlan, Paul
Levasseur, Jeremy
Li, Qingguo
Zhu, Xiaodan
contents Real-world time series come with text: metadata, descriptions, news, reports. Yet time series foundation models process numerical sequences in isolation, and the multimodal text-and-time-series models that attempt to bridge the two all adapt a pretrained language model post hoc, inheriting representations shaped without ever seeing temporal data. These models are also evaluated almost exclusively against other multimodal baselines, not against the strongest unimodal foundation models in either domain, leaving open whether joint training is needed at all. We present Chronicle, a compact 324M-parameter decoder-only transformer trained from scratch on natural language and time series within a single unified architecture. Both modalities share the same transformer blocks, attention mechanism, and residual stream; the bulk of pretraining uses unimodal batches so cross-modal capability emerges purely from shared parameters, with a short alignment stage that interleaves the two. To our knowledge, Chronicle is the first model jointly pretrained on text and time series from scratch, and the first multimodal model evaluated against dedicated foundation models in both domains. It matches Gemma-3-270M-PT on 19 NLU tasks, sets a new bar for frozen-embedding time series classification on 24 UCR/UEA datasets, and produces multimodal forecasts on Time-MMD that beat every supervised fusion baseline, all from a single backbone.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20268
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Chronicle: A Multimodal Foundation Model for Joint Language and Time Series Understanding
Quinlan, Paul
Levasseur, Jeremy
Li, Qingguo
Zhu, Xiaodan
Machine Learning
Artificial Intelligence
Computation and Language
Real-world time series come with text: metadata, descriptions, news, reports. Yet time series foundation models process numerical sequences in isolation, and the multimodal text-and-time-series models that attempt to bridge the two all adapt a pretrained language model post hoc, inheriting representations shaped without ever seeing temporal data. These models are also evaluated almost exclusively against other multimodal baselines, not against the strongest unimodal foundation models in either domain, leaving open whether joint training is needed at all. We present Chronicle, a compact 324M-parameter decoder-only transformer trained from scratch on natural language and time series within a single unified architecture. Both modalities share the same transformer blocks, attention mechanism, and residual stream; the bulk of pretraining uses unimodal batches so cross-modal capability emerges purely from shared parameters, with a short alignment stage that interleaves the two. To our knowledge, Chronicle is the first model jointly pretrained on text and time series from scratch, and the first multimodal model evaluated against dedicated foundation models in both domains. It matches Gemma-3-270M-PT on 19 NLU tasks, sets a new bar for frozen-embedding time series classification on 24 UCR/UEA datasets, and produces multimodal forecasts on Time-MMD that beat every supervised fusion baseline, all from a single backbone.
title Chronicle: A Multimodal Foundation Model for Joint Language and Time Series Understanding
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.20268