Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910224419389440 |
|---|---|
| author | Li, Lin Huang, Jiawei Quan, Qihao Li, Dan Li, Boxin Zhang, Xiao Meng, Erli Feng, Wenjie Lou, Jian Ng, See-Kiong |
| author_facet | Li, Lin Huang, Jiawei Quan, Qihao Li, Dan Li, Boxin Zhang, Xiao Meng, Erli Feng, Wenjie Lou, Jian Ng, See-Kiong |
| contents | In this paper, we propose the first VL$\underline{\textbf{M}}$ $\underline{\textbf{a}}$gentic $\underline{\textbf{r}}$easoning framework for few-$\underline{\textbf{s}}$hot multimodal $\underline{\textbf{T}}$ime $\underline{\textbf{S}}$eries $\underline{\textbf{C}}$lassification ($\textbf{MarsTSC}$), which introduces a self-evolving knowledge bank as a dynamic context iteratively refined via reflective agentic reasoning. The framework comprises three collaborative roles: i) Generator conducts reliable classification via reasoning; ii) Reflector diagnoses the root causes of reasoning errors to yield discriminative insights targeting the temporal features overlooked by Generator; iii) Modifier applies verified updates to the knowledge bank to prevent context collapse. We further introduce a test-time update strategy to enable cautious, continuous knowledge bank refinement to mitigate few-shot bias and distribution shift. Extensive experiments across 12 mainstream time series benchmarks demonstrate that $\textbf{MarsTSC}$ delivers substantial and consistent performance gains across 6 VLM backbones, outperforming both classical and foundation model-based time series baselines under few-shot conditions, while producing interpretable rationales that ground each classification decision in human-readable feature evidence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_09395 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Li, Lin Huang, Jiawei Quan, Qihao Li, Dan Li, Boxin Zhang, Xiao Meng, Erli Feng, Wenjie Lou, Jian Ng, See-Kiong Artificial Intelligence Machine Learning Multiagent Systems Multimedia I.2.0; I.2.4; I.5.4 In this paper, we propose the first VL$\underline{\textbf{M}}$ $\underline{\textbf{a}}$gentic $\underline{\textbf{r}}$easoning framework for few-$\underline{\textbf{s}}$hot multimodal $\underline{\textbf{T}}$ime $\underline{\textbf{S}}$eries $\underline{\textbf{C}}$lassification ($\textbf{MarsTSC}$), which introduces a self-evolving knowledge bank as a dynamic context iteratively refined via reflective agentic reasoning. The framework comprises three collaborative roles: i) Generator conducts reliable classification via reasoning; ii) Reflector diagnoses the root causes of reasoning errors to yield discriminative insights targeting the temporal features overlooked by Generator; iii) Modifier applies verified updates to the knowledge bank to prevent context collapse. We further introduce a test-time update strategy to enable cautious, continuous knowledge bank refinement to mitigate few-shot bias and distribution shift. Extensive experiments across 12 mainstream time series benchmarks demonstrate that $\textbf{MarsTSC}$ delivers substantial and consistent performance gains across 6 VLM backbones, outperforming both classical and foundation model-based time series baselines under few-shot conditions, while producing interpretable rationales that ground each classification decision in human-readable feature evidence. |
| title | Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning |
| topic | Artificial Intelligence Machine Learning Multiagent Systems Multimedia I.2.0; I.2.4; I.5.4 |
| url | https://arxiv.org/abs/2605.09395 |