VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908870934265856 |
|---|---|
| author | Khezresmaeilzadeh, Tina Razmara, Parsa Azizi, Seyedarmin Sadeghi, Mohammad Erfan Potraghloo, Erfan Baghaei |
| author_facet | Khezresmaeilzadeh, Tina Razmara, Parsa Azizi, Seyedarmin Sadeghi, Mohammad Erfan Potraghloo, Erfan Baghaei |
| contents | Stock price prediction remains a complex and high-stakes task in financial analysis, traditionally addressed using statistical models or, more recently, language models. In this work, we introduce VISTA (Vision-Language Inference for Stock Time-series Analysis), a novel, training-free framework that leverages Vision-Language Models (VLMs) for multi-modal stock forecasting. VISTA prompts a VLM with both textual representations of historical stock prices and their corresponding line charts to predict future price values. By combining numerical and visual modalities in a zero-shot setting and using carefully designed chain-of-thought prompts, VISTA captures complementary patterns that unimodal approaches often miss. We benchmark VISTA against standard baselines, including ARIMA and text-only LLM-based prompting methods. Experimental results show that VISTA outperforms these baselines by up to 89.83%, demonstrating the effectiveness of multi-modal inference for stock time-series analysis and highlighting the potential of VLMs in financial forecasting tasks without requiring task-specific training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_18570 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis Khezresmaeilzadeh, Tina Razmara, Parsa Azizi, Seyedarmin Sadeghi, Mohammad Erfan Potraghloo, Erfan Baghaei Machine Learning Stock price prediction remains a complex and high-stakes task in financial analysis, traditionally addressed using statistical models or, more recently, language models. In this work, we introduce VISTA (Vision-Language Inference for Stock Time-series Analysis), a novel, training-free framework that leverages Vision-Language Models (VLMs) for multi-modal stock forecasting. VISTA prompts a VLM with both textual representations of historical stock prices and their corresponding line charts to predict future price values. By combining numerical and visual modalities in a zero-shot setting and using carefully designed chain-of-thought prompts, VISTA captures complementary patterns that unimodal approaches often miss. We benchmark VISTA against standard baselines, including ARIMA and text-only LLM-based prompting methods. Experimental results show that VISTA outperforms these baselines by up to 89.83%, demonstrating the effectiveness of multi-modal inference for stock time-series analysis and highlighting the potential of VLMs in financial forecasting tasks without requiring task-specific training. |
| title | VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2505.18570 |