Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ozolcer, Melik, Bae, Sang Won
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915565264699392
author Ozolcer, Melik
Bae, Sang Won
author_facet Ozolcer, Melik
Bae, Sang Won
contents We study a web-deployed, tool-augmented LLM health coach with real users. In a pilot with seven users (280 rated turns), offline policy evaluation (OPE) over factorized decision heads (Tool/Style) shows that a uniform heavy-tool policy raises average value on logs but harms specific subgroups, most notably low-health-literacy/high-self-efficacy users. A lightweight simulator with hidden archetypes further shows that adding a small early information-gain bonus reliably shortens trait identification and improves goal success and pass@3. Together, these early findings indicate an evaluation-first path to personalization: freeze the generator, learn subgroup-aware decision heads on typed rewards (objective tool outcomes and satisfaction), and always report per-archetype metrics to surface subgroup harms that averages obscure.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users
Ozolcer, Melik
Bae, Sang Won
Artificial Intelligence
Computation and Language
We study a web-deployed, tool-augmented LLM health coach with real users. In a pilot with seven users (280 rated turns), offline policy evaluation (OPE) over factorized decision heads (Tool/Style) shows that a uniform heavy-tool policy raises average value on logs but harms specific subgroups, most notably low-health-literacy/high-self-efficacy users. A lightweight simulator with hidden archetypes further shows that adding a small early information-gain bonus reliably shortens trait identification and improves goal success and pass@3. Together, these early findings indicate an evaluation-first path to personalization: freeze the generator, learn subgroup-aware decision heads on typed rewards (objective tool outcomes and satisfaction), and always report per-archetype metrics to surface subgroup harms that averages obscure.
title Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.17173