Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Albers, Nele, de Groot, Esra Cemre Su, Keijsers, Loes, Hillegers, Manon H., Krahmer, Emiel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912724285390848
author Albers, Nele
de Groot, Esra Cemre Su
Keijsers, Loes
Hillegers, Manon H.
Krahmer, Emiel
author_facet Albers, Nele
de Groot, Esra Cemre Su
Keijsers, Loes
Hillegers, Manon H.
Krahmer, Emiel
contents Personalizing digital applications for health behavior change is a promising route to making them more engaging and effective. This especially holds for approaches that adapt to users and their specific states (e.g., motivation, knowledge, wants) over time. However, developing such approaches requires making many design choices, whose effectiveness is difficult to predict from literature and costly to evaluate in practice. In this work, we explore whether large language models (LLMs) can be used out-of-the-box to generate samples of user interactions that provide useful information for training reinforcement learning models for digital behavior change settings. Using real user data from four large behavior change studies as comparison, we show that LLM-generated samples can be useful in the absence of real data. Comparisons to the samples provided by human raters further show that LLM-generated samples reach the performance of human raters. Additional analyses of different prompting strategies including shorter and longer prompt variants, chain-of-thought prompting, and few-shot prompting show that the relative effectiveness of different strategies depends on both the study and the LLM with also relatively large differences between prompt paraphrases alone. We provide recommendations for how LLM-generated samples can be useful in practice.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
Albers, Nele
de Groot, Esra Cemre Su
Keijsers, Loes
Hillegers, Manon H.
Krahmer, Emiel
Machine Learning
Artificial Intelligence
Human-Computer Interaction
Personalizing digital applications for health behavior change is a promising route to making them more engaging and effective. This especially holds for approaches that adapt to users and their specific states (e.g., motivation, knowledge, wants) over time. However, developing such approaches requires making many design choices, whose effectiveness is difficult to predict from literature and costly to evaluate in practice. In this work, we explore whether large language models (LLMs) can be used out-of-the-box to generate samples of user interactions that provide useful information for training reinforcement learning models for digital behavior change settings. Using real user data from four large behavior change studies as comparison, we show that LLM-generated samples can be useful in the absence of real data. Comparisons to the samples provided by human raters further show that LLM-generated samples reach the performance of human raters. Additional analyses of different prompting strategies including shorter and longer prompt variants, chain-of-thought prompting, and few-shot prompting show that the relative effectiveness of different strategies depends on both the study and the LLM with also relatively large differences between prompt paraphrases alone. We provide recommendations for how LLM-generated samples can be useful in practice.
title Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
topic Machine Learning
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2511.17630