Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Abdulhai, Marwa, Cheng, Ryan, Clay, Donovan, Althoff, Tim, Levine, Sergey, Jaques, Natasha
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915589617876992
author Abdulhai, Marwa
Cheng, Ryan
Clay, Donovan
Althoff, Tim
Levine, Sergey
Jaques, Natasha
author_facet Abdulhai, Marwa
Cheng, Ryan
Clay, Donovan
Althoff, Tim
Levine, Sergey
Jaques, Natasha
contents Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf LLMs often drift from their assigned personas, contradict earlier statements, or abandon role-appropriate behavior. We introduce a unified framework for evaluating and improving persona consistency in LLM-generated dialogue. We define three automatic metrics: prompt-to-line consistency, line-to-line consistency, and Q&A consistency, that capture different types of persona drift and validate each against human annotations. Using these metrics as reward signals, we apply multi-turn reinforcement learning to fine-tune LLMs for three user roles: a patient, a student, and a social chat partner. Our method reduces inconsistency by over 55%, resulting in more coherent and faithful simulated users.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
Abdulhai, Marwa
Cheng, Ryan
Clay, Donovan
Althoff, Tim
Levine, Sergey
Jaques, Natasha
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf LLMs often drift from their assigned personas, contradict earlier statements, or abandon role-appropriate behavior. We introduce a unified framework for evaluating and improving persona consistency in LLM-generated dialogue. We define three automatic metrics: prompt-to-line consistency, line-to-line consistency, and Q&A consistency, that capture different types of persona drift and validate each against human annotations. Using these metrics as reward signals, we apply multi-turn reinforcement learning to fine-tune LLMs for three user roles: a patient, a student, and a social chat partner. Our method reduces inconsistency by over 55%, resulting in more coherent and faithful simulated users.
title Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.00222