Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, ChengAo, Yu, Wenchao, Zhao, Ziming, Song, Dongjin, Cheng, Wei, Chen, Haifeng, Ni, Jingchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914125851918336
author Shen, ChengAo
Yu, Wenchao
Zhao, Ziming
Song, Dongjin
Cheng, Wei
Chen, Haifeng
Ni, Jingchao
author_facet Shen, ChengAo
Yu, Wenchao
Zhao, Ziming
Song, Dongjin
Cheng, Wei
Chen, Haifeng
Ni, Jingchao
contents Time series, typically represented as numerical sequences, can also be transformed into images and texts, offering multi-modal views (MMVs) of the same underlying signal. These MMVs can reveal complementary patterns and enable the use of powerful pre-trained large models, such as large vision models (LVMs), for long-term time series forecasting (LTSF). However, as we identified in this work, the state-of-the-art (SOTA) LVM-based forecaster poses an inductive bias towards "forecasting periods". To harness this bias, we propose DMMV, a novel decomposition-based multi-modal view framework that leverages trend-seasonal decomposition and a novel backcast-residual based adaptive decomposition to integrate MMVs for LTSF. Comparative evaluations against 14 SOTA models across diverse datasets show that DMMV outperforms single-view and existing multi-modal baselines, achieving the best mean squared error (MSE) on 6 out of 8 benchmark datasets. The code for this paper is available at: https://github.com/D2I-Group/dmmv.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24003
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting
Shen, ChengAo
Yu, Wenchao
Zhao, Ziming
Song, Dongjin
Cheng, Wei
Chen, Haifeng
Ni, Jingchao
Machine Learning
Artificial Intelligence
Time series, typically represented as numerical sequences, can also be transformed into images and texts, offering multi-modal views (MMVs) of the same underlying signal. These MMVs can reveal complementary patterns and enable the use of powerful pre-trained large models, such as large vision models (LVMs), for long-term time series forecasting (LTSF). However, as we identified in this work, the state-of-the-art (SOTA) LVM-based forecaster poses an inductive bias towards "forecasting periods". To harness this bias, we propose DMMV, a novel decomposition-based multi-modal view framework that leverages trend-seasonal decomposition and a novel backcast-residual based adaptive decomposition to integrate MMVs for LTSF. Comparative evaluations against 14 SOTA models across diverse datasets show that DMMV outperforms single-view and existing multi-modal baselines, achieving the best mean squared error (MSE) on 6 out of 8 benchmark datasets. The code for this paper is available at: https://github.com/D2I-Group/dmmv.
title Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.24003