Outcome-based Reinforcement Learning to Predict the Future

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Turtel, Benjamin, Franklin, Danny, Skotheim, Kris, Hewitt, Luke, Schoenegger, Philipp
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908682945560576
author Turtel, Benjamin
Franklin, Danny
Skotheim, Kris
Hewitt, Luke
Schoenegger, Philipp
author_facet Turtel, Benjamin
Franklin, Danny
Skotheim, Kris
Hewitt, Luke
Schoenegger, Philipp
contents Reinforcement Learning with Verifiable Rewards (RLVR) has been an effective approach for improving Large Language Models' reasoning in domains such as coding and mathematics. Here, we apply RLVR methods towards forecasting future real-world events - a challenging task for RL due to the very noisy (and delayed) outcomes involved. Using a novel dataset of recent questions from a prediction market, and accompanying relevant news headlines, we show that a compact (14B) reasoning model can be trained to match or surpass the predictive accuracy of frontier models like o1, while greatly improving probabilistic calibration. The model's performance is also practically meaningful: in a Polymarket trading simulation, we estimate that its bets would have yielded a return on investment of over 10% across all questions in the test set. We detail and compare approaches used in training our model, including augmenting our training-data with synthetic prediction questions, guardrails for learning stability, and median prediction sampling at inference-time.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17989
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Outcome-based Reinforcement Learning to Predict the Future
Turtel, Benjamin
Franklin, Danny
Skotheim, Kris
Hewitt, Luke
Schoenegger, Philipp
Machine Learning
Artificial Intelligence
Reinforcement Learning with Verifiable Rewards (RLVR) has been an effective approach for improving Large Language Models' reasoning in domains such as coding and mathematics. Here, we apply RLVR methods towards forecasting future real-world events - a challenging task for RL due to the very noisy (and delayed) outcomes involved. Using a novel dataset of recent questions from a prediction market, and accompanying relevant news headlines, we show that a compact (14B) reasoning model can be trained to match or surpass the predictive accuracy of frontier models like o1, while greatly improving probabilistic calibration. The model's performance is also practically meaningful: in a Polymarket trading simulation, we estimate that its bets would have yielded a return on investment of over 10% across all questions in the test set. We detail and compare approaches used in training our model, including augmenting our training-data with synthetic prediction questions, guardrails for learning stability, and median prediction sampling at inference-time.
title Outcome-based Reinforcement Learning to Predict the Future
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.17989