Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Abdulhai, Marwa, Cheng, Ryan, Clay, Donovan, Althoff, Tim, Levine, Sergey, Jaques, Natasha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
by: Wan, Yanming, et al.
Published: (2025)
by: Wan, Yanming, et al.
Published: (2025)
How LLMs Distort Our Written Language
by: Abdulhai, Marwa, et al.
Published: (2026)
by: Abdulhai, Marwa, et al.
Published: (2026)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
Virtual Personas for Language Models via an Anthology of Backstories
by: Moon, Suhong, et al.
Published: (2024)
by: Moon, Suhong, et al.
Published: (2024)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Maastricht University at AMIYA: Adapting LLMs for Dialectal Arabic using Fine-tuning and MBR Decoding
by: Alali, Abdulhai, et al.
Published: (2026)
by: Alali, Abdulhai, et al.
Published: (2026)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
by: Zhou, Yifei, et al.
Published: (2024)
by: Zhou, Yifei, et al.
Published: (2024)
Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
by: Chen, Xingwu, et al.
Published: (2026)
by: Chen, Xingwu, et al.
Published: (2026)
Infer Human's Intentions Before Following Natural Language Instructions
by: Wan, Yanming, et al.
Published: (2024)
by: Wan, Yanming, et al.
Published: (2024)
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
by: Nam, Hyunji, et al.
Published: (2026)
by: Nam, Hyunji, et al.
Published: (2026)
Consistency of Large Reasoning Models Under Multi-Turn Attacks
by: Li, Yubo, et al.
Published: (2026)
by: Li, Yubo, et al.
Published: (2026)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL
by: Hua, Harper, et al.
Published: (2026)
by: Hua, Harper, et al.
Published: (2026)
MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
by: Lee, Kyungro, et al.
Published: (2025)
by: Lee, Kyungro, et al.
Published: (2025)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
by: Wu, Mian, et al.
Published: (2025)
by: Wu, Mian, et al.
Published: (2025)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
by: Wang, Kangrui, et al.
Published: (2025)
by: Wang, Kangrui, et al.
Published: (2025)
Correcting misinformation on social media with a large language model
by: Zhou, Xinyi, et al.
Published: (2024)
by: Zhou, Xinyi, et al.
Published: (2024)
Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
by: Mehri, Shuhaib, et al.
Published: (2026)
by: Mehri, Shuhaib, et al.
Published: (2026)
Adaptive Stopping for Multi-Turn LLM Reasoning
by: Zhou, Xiaofan, et al.
Published: (2026)
by: Zhou, Xiaofan, et al.
Published: (2026)
Enhancing Consistency of Werewolf AI through Dialogue Summarization and Persona Information
by: Tanaka, Yoshiki, et al.
Published: (2026)
by: Tanaka, Yoshiki, et al.
Published: (2026)
Drift No More? Context Equilibria in Multi-Turn LLM Interactions
by: Dongre, Vardhan, et al.
Published: (2025)
by: Dongre, Vardhan, et al.
Published: (2025)
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
by: Prabhakar, Akshara, et al.
Published: (2025)
by: Prabhakar, Akshara, et al.
Published: (2025)
ReaLJam: Real-Time Human-AI Music Jamming with Reinforcement Learning-Tuned Transformers
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models
by: Li, Yubo, et al.
Published: (2025)
by: Li, Yubo, et al.
Published: (2025)
PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
by: Chai, Huacan, et al.
Published: (2025)
by: Chai, Huacan, et al.
Published: (2025)
The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
by: Cheng, Youyou, et al.
Published: (2026)
by: Cheng, Youyou, et al.
Published: (2026)
PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions
by: Kyung, Daeun, et al.
Published: (2025)
by: Kyung, Daeun, et al.
Published: (2025)
Discourse Diversity in Multi-Turn Empathic Dialogue
by: Zhan, Hongli, et al.
Published: (2026)
by: Zhan, Hongli, et al.
Published: (2026)
The Need for a Socially-Grounded Persona Framework for User Simulation
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
by: Lee, Young-Suk, et al.
Published: (2024)
by: Lee, Young-Suk, et al.
Published: (2024)
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
by: Crouse, Maxwell, et al.
Published: (2026)
by: Crouse, Maxwell, et al.
Published: (2026)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
Similar Items
-
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025) -
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
by: Wan, Yanming, et al.
Published: (2025) -
How LLMs Distort Our Written Language
by: Abdulhai, Marwa, et al.
Published: (2026) -
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
by: Chopra, Harshita, et al.
Published: (2026) -
Virtual Personas for Language Models via an Anthology of Backstories
by: Moon, Suhong, et al.
Published: (2024)