Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Linlu, Sha, Fei, Allen, Kelsey, Kim, Yoon, Linzen, Tal, van Steenkiste, Sjoerd
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911377350721536
author Qiu, Linlu
Sha, Fei
Allen, Kelsey
Kim, Yoon
Linzen, Tal
van Steenkiste, Sjoerd
author_facet Qiu, Linlu
Sha, Fei
Allen, Kelsey
Kim, Yoon
Linzen, Tal
van Steenkiste, Sjoerd
contents Large language models (LLMs) are increasingly used as agents that interact with users and with the world. To do so successfully, LLMs must construct representations of the world and form probabilistic beliefs about them. To provide personalized recommendations, for example, the LLM needs to infer a user's preferences from their behavior over multiple interactions. The Bayesian inference framework lays out the optimal way for an agent to update its beliefs as it receives new information. We first show that LLMs fall far short of the standard defined by the Bayesian framework. We then show that by teaching LLMs to mimic the predictions of the normative Bayesian model, we can dramatically improve their ability to update their beliefs; this ability generalizes to new tasks. We conclude that LLMs can effectively learn reasoning skills from examples and generalize those skills to new domains.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17523
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
Qiu, Linlu
Sha, Fei
Allen, Kelsey
Kim, Yoon
Linzen, Tal
van Steenkiste, Sjoerd
Computation and Language
Artificial Intelligence
Large language models (LLMs) are increasingly used as agents that interact with users and with the world. To do so successfully, LLMs must construct representations of the world and form probabilistic beliefs about them. To provide personalized recommendations, for example, the LLM needs to infer a user's preferences from their behavior over multiple interactions. The Bayesian inference framework lays out the optimal way for an agent to update its beliefs as it receives new information. We first show that LLMs fall far short of the standard defined by the Bayesian framework. We then show that by teaching LLMs to mimic the predictions of the normative Bayesian model, we can dramatically improve their ability to update their beliefs; this ability generalizes to new tasks. We conclude that LLMs can effectively learn reasoning skills from examples and generalize those skills to new domains.
title Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.17523