TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913119490539520 |
|---|---|
| author | Zumarraga, Nicolas Kaar, Thomas Wang, Ning Tennien, William Hasanli, Alpay Rosenblattl, Max Wu, Fan Riehl, Kevin Xu, Maxwell A. Kreft, Markus O'Sullivan, Kevin Fleisch, Elgar Schmiedmayer, Paul Jakob, Robert Langer, Patrick |
| author_facet | Zumarraga, Nicolas Kaar, Thomas Wang, Ning Tennien, William Hasanli, Alpay Rosenblattl, Max Wu, Fan Riehl, Kevin Xu, Maxwell A. Kreft, Markus O'Sullivan, Kevin Fleisch, Elgar Schmiedmayer, Paul Jakob, Robert Langer, Patrick |
| contents | Time Series Language Models (TSLMs) promise reasoning over real-world temporal data, but their ability to retrieve and reason over long time-series remains largely untested. We introduce TS-Haystack, a multi-domain retrieval benchmark with ten event-grounded question-answering tasks over contexts from 100 seconds to 24 hours, spanning direct retrieval, temporal reasoning, multi-step reasoning, and contextual anomaly detection. Existing TSLMs exhibit severe long-context degradation: accuracy declines with context length, direct-tokenization models run out of memory beyond 100 seconds on high-rate signals, and time-interval-grounded tasks collapse toward near-zero accuracy when increasing the time-series lengths, aligning with existing literature on text and multi-modal long context retrieval. An agentic retrieval framework using specialized time-series classifier tools matches or outperforms SoTA TSLMs on 9 of 10 tasks, highlighting agentic retrieval as a promising approach for long-context TSLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_14200 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning Zumarraga, Nicolas Kaar, Thomas Wang, Ning Tennien, William Hasanli, Alpay Rosenblattl, Max Wu, Fan Riehl, Kevin Xu, Maxwell A. Kreft, Markus O'Sullivan, Kevin Fleisch, Elgar Schmiedmayer, Paul Jakob, Robert Langer, Patrick Machine Learning Time Series Language Models (TSLMs) promise reasoning over real-world temporal data, but their ability to retrieve and reason over long time-series remains largely untested. We introduce TS-Haystack, a multi-domain retrieval benchmark with ten event-grounded question-answering tasks over contexts from 100 seconds to 24 hours, spanning direct retrieval, temporal reasoning, multi-step reasoning, and contextual anomaly detection. Existing TSLMs exhibit severe long-context degradation: accuracy declines with context length, direct-tokenization models run out of memory beyond 100 seconds on high-rate signals, and time-interval-grounded tasks collapse toward near-zero accuracy when increasing the time-series lengths, aligning with existing literature on text and multi-modal long context retrieval. An agentic retrieval framework using specialized time-series classifier tools matches or outperforms SoTA TSLMs on 9 of 10 tasks, highlighting agentic retrieval as a promising approach for long-context TSLMs. |
| title | TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2602.14200 |