ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karger, Ezra, Bastani, Houtan, Yueh-Han, Chen, Jacobs, Zachary, Halawi, Danny, Zhang, Fred, Tetlock, Philip E.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909515965792256
author Karger, Ezra
Bastani, Houtan
Yueh-Han, Chen
Jacobs, Zachary
Halawi, Danny
Zhang, Fred
Tetlock, Philip E.
author_facet Karger, Ezra
Bastani, Houtan
Yueh-Han, Chen
Jacobs, Zachary
Halawi, Danny
Zhang, Fred
Tetlock, Philip E.
contents Forecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark ($N=200$). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM ($p$-value $<0.001$). We display system and human scores in a public leaderboard at www.forecastbench.org.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
Karger, Ezra
Bastani, Houtan
Yueh-Han, Chen
Jacobs, Zachary
Halawi, Danny
Zhang, Fred
Tetlock, Philip E.
Machine Learning
Artificial Intelligence
Computation and Language
Forecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark ($N=200$). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM ($p$-value $<0.001$). We display system and human scores in a public leaderboard at www.forecastbench.org.
title ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2409.19839