Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Chopra, Harshita, Ghate, Kshitish, Caliskan, Aylin, Kohno, Tadayoshi, Shah, Chirag, Jaques, Natasha |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Reasoning in General Purpose Agents via Structured Meta-Cognition
by: Light, Dean, et al.
Published: (2026)
by: Light, Dean, et al.
Published: (2026)
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
Feedback-Aware Monte Carlo Tree Search for Efficient Information Seeking in Goal-Oriented Conversations
by: Chopra, Harshita, et al.
Published: (2025)
by: Chopra, Harshita, et al.
Published: (2025)
Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models
by: Veerendranath, Vishruth, et al.
Published: (2024)
by: Veerendranath, Vishruth, et al.
Published: (2024)
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026)
by: Subramani, Nishant, et al.
Published: (2026)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
by: Iqbal, Umar, et al.
Published: (2023)
by: Iqbal, Umar, et al.
Published: (2023)
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
by: Dash, Saloni, et al.
Published: (2025)
by: Dash, Saloni, et al.
Published: (2025)
Generative Value Conflicts Reveal LLM Priorities
by: Liu, Andy, et al.
Published: (2025)
by: Liu, Andy, et al.
Published: (2025)
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
ChatGPT Perpetuates Gender Bias in Machine Translation and Ignores Non-Gendered Pronouns: Findings across Bengali and Five other Low-Resource Languages
by: Ghosh, Sourojit, et al.
Published: (2023)
by: Ghosh, Sourojit, et al.
Published: (2023)
A Taxonomy of Stereotype Content in Large Language Models
by: Nicolas, Gandalf, et al.
Published: (2024)
by: Nicolas, Gandalf, et al.
Published: (2024)
REALM: A Dataset of Real-World LLM Use Cases
by: Cheng, Jingwen, et al.
Published: (2025)
by: Cheng, Jingwen, et al.
Published: (2025)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
by: Wilson, Kyra, et al.
Published: (2024)
by: Wilson, Kyra, et al.
Published: (2024)
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
by: Nam, Hyunji, et al.
Published: (2026)
by: Nam, Hyunji, et al.
Published: (2026)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
by: Lee, Huije, et al.
Published: (2026)
by: Lee, Huije, et al.
Published: (2026)
Developing Story: Case Studies of Generative AI's Use in Journalism
by: Brigham, Natalie Grace, et al.
Published: (2024)
by: Brigham, Natalie Grace, et al.
Published: (2024)
Beyond One-Way Influence: Bidirectional Opinion Dynamics in Multi-Turn Human-LLM Interactions
by: Jiang, Yuyang, et al.
Published: (2025)
by: Jiang, Yuyang, et al.
Published: (2025)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
by: Bai, Yuqi, et al.
Published: (2025)
by: Bai, Yuqi, et al.
Published: (2025)
PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions
by: Kyung, Daeun, et al.
Published: (2025)
by: Kyung, Daeun, et al.
Published: (2025)
Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
by: Wang, Shuting, et al.
Published: (2025)
by: Wang, Shuting, et al.
Published: (2025)
Point of Order: Action-Aware LLM Persona Modeling for Realistic Civic Simulation
by: Merrill, Scott, et al.
Published: (2025)
by: Merrill, Scott, et al.
Published: (2025)
Subjective Behaviors and Preferences in LLM: Language of Browsing
by: Sundaresan, Sai, et al.
Published: (2025)
by: Sundaresan, Sai, et al.
Published: (2025)
PersonaGym: Evaluating Persona Agents and LLMs
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
Learning to Cooperate with Humans using Generative Agents
by: Liang, Yancheng, et al.
Published: (2024)
by: Liang, Yancheng, et al.
Published: (2024)
$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy
by: Wilson, Kyra, et al.
Published: (2025)
by: Wilson, Kyra, et al.
Published: (2025)
Population-Aligned Persona Generation for LLM-based Social Simulation
by: Hu, Zhengyu, et al.
Published: (2025)
by: Hu, Zhengyu, et al.
Published: (2025)
Simulating Misinformation Vulnerabilities With Agent Personas
by: Farr, David, et al.
Published: (2025)
by: Farr, David, et al.
Published: (2025)
PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents
by: Wang, Minjia, et al.
Published: (2026)
by: Wang, Minjia, et al.
Published: (2026)
The Need for a Socially-Grounded Persona Framework for User Simulation
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
Chasing Random: Instruction Selection Strategies Fail to Generalize
by: Diddee, Harshita, et al.
Published: (2024)
by: Diddee, Harshita, et al.
Published: (2024)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
by: Wang, Xintao, et al.
Published: (2025)
by: Wang, Xintao, et al.
Published: (2025)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
by: Kim, Serin, et al.
Published: (2026)
by: Kim, Serin, et al.
Published: (2026)
Similar Items
-
Deep Reasoning in General Purpose Agents via Structured Meta-Cognition
by: Light, Dean, et al.
Published: (2026) -
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
by: Ghate, Kshitish, et al.
Published: (2025) -
Feedback-Aware Monte Carlo Tree Search for Efficient Information Seeking in Goal-Oriented Conversations
by: Chopra, Harshita, et al.
Published: (2025) -
Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models
by: Veerendranath, Vishruth, et al.
Published: (2024) -
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
by: Ghate, Kshitish, et al.
Published: (2025)