VIFO: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yanlong, Yu, Hang, Xu, Jian, Ma, Fei, Zhang, Hongkang, Feng, Tongtong, Zhang, Zijian, Huang, Shao-Lun, Sun, Danny Dongning, Zhang, Xiao-Ping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914072870518784
author Wang, Yanlong
Yu, Hang
Xu, Jian
Ma, Fei
Zhang, Hongkang
Feng, Tongtong
Zhang, Zijian
Huang, Shao-Lun
Sun, Danny Dongning
Zhang, Xiao-Ping
author_facet Wang, Yanlong
Yu, Hang
Xu, Jian
Ma, Fei
Zhang, Hongkang
Feng, Tongtong
Zhang, Zijian
Huang, Shao-Lun
Sun, Danny Dongning
Zhang, Xiao-Ping
contents Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores crucial cross-channel dependencies. Concurrently, existing multimodal approaches have not fully exploited the power of large vision models (LVMs) to interpret spatiotemporal data. Additionally, there remains significant unexplored potential in leveraging the advantages of information extraction from different modalities to enhance time series forecasting performance. To address these gaps, we propose the VIFO, a cross-modal forecasting model. VIFO uniquely renders multivariate time series into image, enabling pre-trained LVM to extract complex cross-channel patterns that are invisible to channel-independent models. These visual features are then aligned and fused with representations from the time series modality. By freezing the LVM and training only 7.45% of its parameters, VIFO achieves competitive performance on multiple benchmarks, offering an efficient and effective solution for capturing cross-variable relationships in
format Preprint
id arxiv_https___arxiv_org_abs_2510_03244
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VIFO: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion
Wang, Yanlong
Yu, Hang
Xu, Jian
Ma, Fei
Zhang, Hongkang
Feng, Tongtong
Zhang, Zijian
Huang, Shao-Lun
Sun, Danny Dongning
Zhang, Xiao-Ping
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores crucial cross-channel dependencies. Concurrently, existing multimodal approaches have not fully exploited the power of large vision models (LVMs) to interpret spatiotemporal data. Additionally, there remains significant unexplored potential in leveraging the advantages of information extraction from different modalities to enhance time series forecasting performance. To address these gaps, we propose the VIFO, a cross-modal forecasting model. VIFO uniquely renders multivariate time series into image, enabling pre-trained LVM to extract complex cross-channel patterns that are invisible to channel-independent models. These visual features are then aligned and fused with representations from the time series modality. By freezing the LVM and training only 7.45% of its parameters, VIFO achieves competitive performance on multiple benchmarks, offering an efficient and effective solution for capturing cross-variable relationships in
title VIFO: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.03244