Prompting Underestimates LLM Capability for Time Series Classification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Schumacher, Dan, Nourbakhsh, Erfan, Slavin, Rocky, Rios, Anthony
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911508146946048
author Schumacher, Dan
Nourbakhsh, Erfan
Slavin, Rocky
Rios, Anthony
author_facet Schumacher, Dan
Nourbakhsh, Erfan
Slavin, Rocky
Rios, Anthony
contents Prompt-based evaluations suggest that large language models (LLMs) perform poorly on time series classification, raising doubts about whether they encode meaningful temporal structure. We show that this conclusion reflects limitations of prompt-based generation rather than the model's representational capacity by directly comparing prompt outputs with linear probes over the same internal representations. While zero-shot prompting performs near chance, linear probes improve average F1 from 0.15-0.26 to 0.61-0.67, often matching or exceeding specialized time series models. Layer-wise analyses further show that class-discriminative time series information emerges in early transformer layers and is amplified by visual and multimodal inputs. Together, these results demonstrate a systematic mismatch between what LLMs internally represent and what prompt-based evaluation reveals, leading current evaluations to underestimate their time series understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03464
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Prompting Underestimates LLM Capability for Time Series Classification
Schumacher, Dan
Nourbakhsh, Erfan
Slavin, Rocky
Rios, Anthony
Computation and Language
Prompt-based evaluations suggest that large language models (LLMs) perform poorly on time series classification, raising doubts about whether they encode meaningful temporal structure. We show that this conclusion reflects limitations of prompt-based generation rather than the model's representational capacity by directly comparing prompt outputs with linear probes over the same internal representations. While zero-shot prompting performs near chance, linear probes improve average F1 from 0.15-0.26 to 0.61-0.67, often matching or exceeding specialized time series models. Layer-wise analyses further show that class-discriminative time series information emerges in early transformer layers and is amplified by visual and multimodal inputs. Together, these results demonstrate a systematic mismatch between what LLMs internally represent and what prompt-based evaluation reveals, leading current evaluations to underestimate their time series understanding.
title Prompting Underestimates LLM Capability for Time Series Classification
topic Computation and Language
url https://arxiv.org/abs/2601.03464