TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zumarraga, Nicolas, Kaar, Thomas, Wang, Ning, Tennien, William, Hasanli, Alpay, Rosenblattl, Max, Wu, Fan, Riehl, Kevin, Xu, Maxwell A., Kreft, Markus, O'Sullivan, Kevin, Fleisch, Elgar, Schmiedmayer, Paul, Jakob, Robert, Langer, Patrick
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913119490539520
author Zumarraga, Nicolas
Kaar, Thomas
Wang, Ning
Tennien, William
Hasanli, Alpay
Rosenblattl, Max
Wu, Fan
Riehl, Kevin
Xu, Maxwell A.
Kreft, Markus
O'Sullivan, Kevin
Fleisch, Elgar
Schmiedmayer, Paul
Jakob, Robert
Langer, Patrick
author_facet Zumarraga, Nicolas
Kaar, Thomas
Wang, Ning
Tennien, William
Hasanli, Alpay
Rosenblattl, Max
Wu, Fan
Riehl, Kevin
Xu, Maxwell A.
Kreft, Markus
O'Sullivan, Kevin
Fleisch, Elgar
Schmiedmayer, Paul
Jakob, Robert
Langer, Patrick
contents Time Series Language Models (TSLMs) promise reasoning over real-world temporal data, but their ability to retrieve and reason over long time-series remains largely untested. We introduce TS-Haystack, a multi-domain retrieval benchmark with ten event-grounded question-answering tasks over contexts from 100 seconds to 24 hours, spanning direct retrieval, temporal reasoning, multi-step reasoning, and contextual anomaly detection. Existing TSLMs exhibit severe long-context degradation: accuracy declines with context length, direct-tokenization models run out of memory beyond 100 seconds on high-rate signals, and time-interval-grounded tasks collapse toward near-zero accuracy when increasing the time-series lengths, aligning with existing literature on text and multi-modal long context retrieval. An agentic retrieval framework using specialized time-series classifier tools matches or outperforms SoTA TSLMs on 9 of 10 tasks, highlighting agentic retrieval as a promising approach for long-context TSLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14200
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning
Zumarraga, Nicolas
Kaar, Thomas
Wang, Ning
Tennien, William
Hasanli, Alpay
Rosenblattl, Max
Wu, Fan
Riehl, Kevin
Xu, Maxwell A.
Kreft, Markus
O'Sullivan, Kevin
Fleisch, Elgar
Schmiedmayer, Paul
Jakob, Robert
Langer, Patrick
Machine Learning
Time Series Language Models (TSLMs) promise reasoning over real-world temporal data, but their ability to retrieve and reason over long time-series remains largely untested. We introduce TS-Haystack, a multi-domain retrieval benchmark with ten event-grounded question-answering tasks over contexts from 100 seconds to 24 hours, spanning direct retrieval, temporal reasoning, multi-step reasoning, and contextual anomaly detection. Existing TSLMs exhibit severe long-context degradation: accuracy declines with context length, direct-tokenization models run out of memory beyond 100 seconds on high-rate signals, and time-interval-grounded tasks collapse toward near-zero accuracy when increasing the time-series lengths, aligning with existing literature on text and multi-modal long context retrieval. An agentic retrieval framework using specialized time-series classifier tools matches or outperforms SoTA TSLMs on 9 of 10 tasks, highlighting agentic retrieval as a promising approach for long-context TSLMs.
title TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning
topic Machine Learning
url https://arxiv.org/abs/2602.14200