LEAF: A Living Benchmark for Event-Augmented Forecasting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Mingtian, Parmar, Mihir, Goyal, Palash, Li, Chun-Liang, Peng, Nanyun, Hartvigsen, Thomas, Yoon, Jinsung, Pfister, Tomas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913133563478016
author Tan, Mingtian
Parmar, Mihir
Goyal, Palash
Li, Chun-Liang
Peng, Nanyun
Hartvigsen, Thomas
Yoon, Jinsung
Pfister, Tomas
author_facet Tan, Mingtian
Parmar, Mihir
Goyal, Palash
Li, Chun-Liang
Peng, Nanyun
Hartvigsen, Thomas
Yoon, Jinsung
Pfister, Tomas
contents Large Language Models (LLMs) are increasingly applied to forecasting. To evaluate this capability while mitigating pre-training data contamination, several living benchmarks have been proposed. However, existing benchmarks either lack the multidimensional events essential for accurate forecasting due to data scarcity, or focus on relatively closed environments. To assess the predictive capabilities of LLMs in complex, real-world scenarios, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including future event probabilities, trend and time series forecasting. LEAF utilizes a recursive retrieval agent system paired with dual-agent cross-validation to provide comprehensive and relevant auxiliary text for forecasting. Evaluating state-of-the-art proprietary and open-weight LLMs, we find that these models can leverage signals extracted from complex events to enhance predictive performance. In the stock domain, we find that LLMs achieve better performance on equities they confidently identify as more predictable. Furthermore, the events demonstrate a strong correlation with the target equities. To this end, LEAF provides a necessary, dynamically updating testbed to continuously track and drive progress in event-driven forecasting tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16358
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LEAF: A Living Benchmark for Event-Augmented Forecasting
Tan, Mingtian
Parmar, Mihir
Goyal, Palash
Li, Chun-Liang
Peng, Nanyun
Hartvigsen, Thomas
Yoon, Jinsung
Pfister, Tomas
Machine Learning
Artificial Intelligence
Large Language Models (LLMs) are increasingly applied to forecasting. To evaluate this capability while mitigating pre-training data contamination, several living benchmarks have been proposed. However, existing benchmarks either lack the multidimensional events essential for accurate forecasting due to data scarcity, or focus on relatively closed environments. To assess the predictive capabilities of LLMs in complex, real-world scenarios, we propose LEAF, the first living benchmark for event-augmented forecasting tasks, including future event probabilities, trend and time series forecasting. LEAF utilizes a recursive retrieval agent system paired with dual-agent cross-validation to provide comprehensive and relevant auxiliary text for forecasting. Evaluating state-of-the-art proprietary and open-weight LLMs, we find that these models can leverage signals extracted from complex events to enhance predictive performance. In the stock domain, we find that LLMs achieve better performance on equities they confidently identify as more predictable. Furthermore, the events demonstrate a strong correlation with the target equities. To this end, LEAF provides a necessary, dynamically updating testbed to continuously track and drive progress in event-driven forecasting tasks.
title LEAF: A Living Benchmark for Event-Augmented Forecasting
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.16358