Can Language Models Use Forecasting Strategies?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pratt, Sarah, Blumberg, Seth, Carolino, Pietro Kreitlon, Morris, Meredith Ringel
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909218657796096
author Pratt, Sarah
Blumberg, Seth
Carolino, Pietro Kreitlon
Morris, Meredith Ringel
author_facet Pratt, Sarah
Blumberg, Seth
Carolino, Pietro Kreitlon
Morris, Meredith Ringel
contents Advances in deep learning systems have allowed large models to match or surpass human accuracy on a number of skills such as image classification, basic programming, and standardized test taking. As the performance of the most capable models begin to saturate on tasks where humans already achieve high accuracy, it becomes necessary to benchmark models on increasingly complex abilities. One such task is forecasting the future outcome of events. In this work we describe experiments using a novel dataset of real world events and associated human predictions, an evaluation metric to measure forecasting ability, and the accuracy of a number of different LLM based forecasting designs on the provided dataset. Additionally, we analyze the performance of the LLM forecasters against human predictions and find that models still struggle to make accurate predictions about the future. Our follow-up experiments indicate this is likely due to models' tendency to guess that most events are unlikely to occur (which tends to be true for many prediction datasets, but does not reflect actual forecasting abilities). We reflect on next steps for developing a systematic and reliable approach to studying LLM forecasting.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04446
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Language Models Use Forecasting Strategies?
Pratt, Sarah
Blumberg, Seth
Carolino, Pietro Kreitlon
Morris, Meredith Ringel
Machine Learning
Artificial Intelligence
Advances in deep learning systems have allowed large models to match or surpass human accuracy on a number of skills such as image classification, basic programming, and standardized test taking. As the performance of the most capable models begin to saturate on tasks where humans already achieve high accuracy, it becomes necessary to benchmark models on increasingly complex abilities. One such task is forecasting the future outcome of events. In this work we describe experiments using a novel dataset of real world events and associated human predictions, an evaluation metric to measure forecasting ability, and the accuracy of a number of different LLM based forecasting designs on the provided dataset. Additionally, we analyze the performance of the LLM forecasters against human predictions and find that models still struggle to make accurate predictions about the future. Our follow-up experiments indicate this is likely due to models' tendency to guess that most events are unlikely to occur (which tends to be true for many prediction datasets, but does not reflect actual forecasting abilities). We reflect on next steps for developing a systematic and reliable approach to studying LLM forecasting.
title Can Language Models Use Forecasting Strategies?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.04446