Scaling Open-Ended Reasoning to Predict the Future

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chandak, Nikhil, Goel, Shashwat, Prabhu, Ameya, Hardt, Moritz, Geiping, Jonas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909981576527872
author Chandak, Nikhil
Goel, Shashwat
Prabhu, Ameya
Hardt, Moritz
Geiping, Jonas
author_facet Chandak, Nikhil
Goel, Shashwat
Prabhu, Ameya
Hardt, Moritz
Geiping, Jonas
contents High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. To scale up training data, we synthesize novel forecasting questions from global events reported in daily news, using a fully automated, careful curation recipe. We train the Qwen3 thinking models on our dataset, OpenForesight. To prevent leakage of future information during training and evaluation, we use an offline news corpus, both for data generation and retrieval in our forecasting system. Guided by a small validation set, we show the benefits of retrieval, and an improved reward function for reinforcement learning (RL). Once we obtain our final forecasting system, we perform held-out testing between May to August 2025. Our specialized model, OpenForecaster 8B, matches much larger proprietary models, with our training improving the accuracy, calibration, and consistency of predictions. We find calibration improvements from forecasting training generalize across popular benchmarks. We open-source all our models, code, and data to make research on language model forecasting broadly accessible.
format Preprint
id arxiv_https___arxiv_org_abs_2512_25070
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Open-Ended Reasoning to Predict the Future
Chandak, Nikhil
Goel, Shashwat
Prabhu, Ameya
Hardt, Moritz
Geiping, Jonas
Machine Learning
Computation and Language
High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. To scale up training data, we synthesize novel forecasting questions from global events reported in daily news, using a fully automated, careful curation recipe. We train the Qwen3 thinking models on our dataset, OpenForesight. To prevent leakage of future information during training and evaluation, we use an offline news corpus, both for data generation and retrieval in our forecasting system. Guided by a small validation set, we show the benefits of retrieval, and an improved reward function for reinforcement learning (RL). Once we obtain our final forecasting system, we perform held-out testing between May to August 2025. Our specialized model, OpenForecaster 8B, matches much larger proprietary models, with our training improving the accuracy, calibration, and consistency of predictions. We find calibration improvements from forecasting training generalize across popular benchmarks. We open-source all our models, code, and data to make research on language model forecasting broadly accessible.
title Scaling Open-Ended Reasoning to Predict the Future
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2512.25070