Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jian-Qiao, Xie, Hanbo, Arumugam, Dilip, Wilson, Robert C., Griffiths, Thomas L.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910006989815808
author Zhu, Jian-Qiao
Xie, Hanbo
Arumugam, Dilip
Wilson, Robert C.
Griffiths, Thomas L.
author_facet Zhu, Jian-Qiao
Xie, Hanbo
Arumugam, Dilip
Wilson, Robert C.
Griffiths, Thomas L.
contents A central goal of cognitive modeling is to develop models that not only predict human behavior but also provide insight into the underlying cognitive mechanisms. While neural network models trained on large-scale behavioral data often achieve strong predictive performance, they typically fall short in offering interpretable explanations of the cognitive processes they capture. In this work, we explore the potential of pretrained large language models (LLMs) to serve as dual-purpose cognitive models--capable of both accurate prediction and interpretable explanation in natural language. Specifically, we employ reinforcement learning with outcome-based rewards to guide LLMs toward generating explicit reasoning traces for explaining human risky choices. Our findings demonstrate that this approach produces high-quality explanations alongside strong quantitative predictions of human decisions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
Zhu, Jian-Qiao
Xie, Hanbo
Arumugam, Dilip
Wilson, Robert C.
Griffiths, Thomas L.
Artificial Intelligence
Computation and Language
A central goal of cognitive modeling is to develop models that not only predict human behavior but also provide insight into the underlying cognitive mechanisms. While neural network models trained on large-scale behavioral data often achieve strong predictive performance, they typically fall short in offering interpretable explanations of the cognitive processes they capture. In this work, we explore the potential of pretrained large language models (LLMs) to serve as dual-purpose cognitive models--capable of both accurate prediction and interpretable explanation in natural language. Specifically, we employ reinforcement learning with outcome-based rewards to guide LLMs toward generating explicit reasoning traces for explaining human risky choices. Our findings demonstrate that this approach produces high-quality explanations alongside strong quantitative predictions of human decisions.
title Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.11614