Saved in:
Bibliographic Details
Main Authors: Prithyani, Vinay, Mohammed, Mohsin, Gadgil, Richa, Buitrago, Ricardo, Jain, Vinija, Chadha, Aman
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2412.17304
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910788922376192
author Prithyani, Vinay
Mohammed, Mohsin
Gadgil, Richa
Buitrago, Ricardo
Jain, Vinija
Chadha, Aman
author_facet Prithyani, Vinay
Mohammed, Mohsin
Gadgil, Richa
Buitrago, Ricardo
Jain, Vinija
Chadha, Aman
contents We build upon time-series classification by leveraging the capabilities of Vision Language Models (VLMs). We find that VLMs produce competitive results after two or less epochs of fine-tuning. We develop a novel approach that incorporates graphical data representations as images in conjunction with numerical data. This approach is rooted in the hypothesis that graphical representations can provide additional contextual information that numerical data alone may not capture. Additionally, providing a graphical representation can circumvent issues such as limited context length faced by LLMs. To further advance this work, we implemented a scalable end-to-end pipeline for training on different scenarios, allowing us to isolate the most effective strategies for transferring learning capabilities from LLMs to Time Series Classification (TSC) tasks. Our approach works with univariate and multivariate time-series data. In addition, we conduct extensive and practical experiments to show how this approach works for time-series classification and generative labels.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17304
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Feasibility of Vision-Language Models for Time-Series Classification
Prithyani, Vinay
Mohammed, Mohsin
Gadgil, Richa
Buitrago, Ricardo
Jain, Vinija
Chadha, Aman
Artificial Intelligence
We build upon time-series classification by leveraging the capabilities of Vision Language Models (VLMs). We find that VLMs produce competitive results after two or less epochs of fine-tuning. We develop a novel approach that incorporates graphical data representations as images in conjunction with numerical data. This approach is rooted in the hypothesis that graphical representations can provide additional contextual information that numerical data alone may not capture. Additionally, providing a graphical representation can circumvent issues such as limited context length faced by LLMs. To further advance this work, we implemented a scalable end-to-end pipeline for training on different scenarios, allowing us to isolate the most effective strategies for transferring learning capabilities from LLMs to Time Series Classification (TSC) tasks. Our approach works with univariate and multivariate time-series data. In addition, we conduct extensive and practical experiments to show how this approach works for time-series classification and generative labels.
title On the Feasibility of Vision-Language Models for Time-Series Classification
topic Artificial Intelligence
url https://arxiv.org/abs/2412.17304