Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914125851918336 |
|---|---|
| author | Shen, ChengAo Yu, Wenchao Zhao, Ziming Song, Dongjin Cheng, Wei Chen, Haifeng Ni, Jingchao |
| author_facet | Shen, ChengAo Yu, Wenchao Zhao, Ziming Song, Dongjin Cheng, Wei Chen, Haifeng Ni, Jingchao |
| contents | Time series, typically represented as numerical sequences, can also be transformed into images and texts, offering multi-modal views (MMVs) of the same underlying signal. These MMVs can reveal complementary patterns and enable the use of powerful pre-trained large models, such as large vision models (LVMs), for long-term time series forecasting (LTSF). However, as we identified in this work, the state-of-the-art (SOTA) LVM-based forecaster poses an inductive bias towards "forecasting periods". To harness this bias, we propose DMMV, a novel decomposition-based multi-modal view framework that leverages trend-seasonal decomposition and a novel backcast-residual based adaptive decomposition to integrate MMVs for LTSF. Comparative evaluations against 14 SOTA models across diverse datasets show that DMMV outperforms single-view and existing multi-modal baselines, achieving the best mean squared error (MSE) on 6 out of 8 benchmark datasets. The code for this paper is available at: https://github.com/D2I-Group/dmmv. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_24003 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting Shen, ChengAo Yu, Wenchao Zhao, Ziming Song, Dongjin Cheng, Wei Chen, Haifeng Ni, Jingchao Machine Learning Artificial Intelligence Time series, typically represented as numerical sequences, can also be transformed into images and texts, offering multi-modal views (MMVs) of the same underlying signal. These MMVs can reveal complementary patterns and enable the use of powerful pre-trained large models, such as large vision models (LVMs), for long-term time series forecasting (LTSF). However, as we identified in this work, the state-of-the-art (SOTA) LVM-based forecaster poses an inductive bias towards "forecasting periods". To harness this bias, we propose DMMV, a novel decomposition-based multi-modal view framework that leverages trend-seasonal decomposition and a novel backcast-residual based adaptive decomposition to integrate MMVs for LTSF. Comparative evaluations against 14 SOTA models across diverse datasets show that DMMV outperforms single-view and existing multi-modal baselines, achieving the best mean squared error (MSE) on 6 out of 8 benchmark datasets. The code for this paper is available at: https://github.com/D2I-Group/dmmv. |
| title | Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2505.24003 |