OccamVTS: Distilling Vision Models to 1% Parameters for Time Series Forecasting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lyu, Sisuo, Zhong, Siru, Ruan, Weilin, Liu, Qingxiang, Wen, Qingsong, Xiong, Hui, Liang, Yuxuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908652190826496
author Lyu, Sisuo
Zhong, Siru
Ruan, Weilin
Liu, Qingxiang
Wen, Qingsong
Xiong, Hui
Liang, Yuxuan
author_facet Lyu, Sisuo
Zhong, Siru
Ruan, Weilin
Liu, Qingxiang
Wen, Qingsong
Xiong, Hui
Liang, Yuxuan
contents Time series forecasting is fundamental to diverse applications, with recent approaches leverage large vision models (LVMs) to capture temporal patterns through visual representations. We reveal that while vision models enhance forecasting performance, 99% of their parameters are unnecessary for time series tasks. Through cross-modal analysis, we find that time series align with low-level textural features but not high-level semantics, which can impair forecasting accuracy. We propose OccamVTS, a knowledge distillation framework that extracts only the essential 1% of predictive information from LVMs into lightweight networks. Using pre-trained LVMs as privileged teachers, OccamVTS employs pyramid-style feature alignment combined with correlation and feature distillation to transfer beneficial patterns while filtering out semantic noise. Counterintuitively, this aggressive parameter reduction improves accuracy by eliminating overfitting to irrelevant visual features while preserving essential temporal patterns. Extensive experiments across multiple benchmark datasets demonstrate that OccamVTS consistently achieves state-of-the-art performance with only 1% of the original parameters, particularly excelling in few-shot and zero-shot scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01727
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OccamVTS: Distilling Vision Models to 1% Parameters for Time Series Forecasting
Lyu, Sisuo
Zhong, Siru
Ruan, Weilin
Liu, Qingxiang
Wen, Qingsong
Xiong, Hui
Liang, Yuxuan
Machine Learning
Computer Vision and Pattern Recognition
Time series forecasting is fundamental to diverse applications, with recent approaches leverage large vision models (LVMs) to capture temporal patterns through visual representations. We reveal that while vision models enhance forecasting performance, 99% of their parameters are unnecessary for time series tasks. Through cross-modal analysis, we find that time series align with low-level textural features but not high-level semantics, which can impair forecasting accuracy. We propose OccamVTS, a knowledge distillation framework that extracts only the essential 1% of predictive information from LVMs into lightweight networks. Using pre-trained LVMs as privileged teachers, OccamVTS employs pyramid-style feature alignment combined with correlation and feature distillation to transfer beneficial patterns while filtering out semantic noise. Counterintuitively, this aggressive parameter reduction improves accuracy by eliminating overfitting to irrelevant visual features while preserving essential temporal patterns. Extensive experiments across multiple benchmark datasets demonstrate that OccamVTS consistently achieves state-of-the-art performance with only 1% of the original parameters, particularly excelling in few-shot and zero-shot scenarios.
title OccamVTS: Distilling Vision Models to 1% Parameters for Time Series Forecasting
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.01727