PersonaGym: Evaluating Persona Agents and LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Samuel, Vinay, Zou, Henry Peng, Zhou, Yue, Chaudhari, Shreyas, Kalyan, Ashwin, Rajpurohit, Tanmay, Deshpande, Ameet, Narasimhan, Karthik, Murahari, Vishvak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909772022808576
author Samuel, Vinay
Zou, Henry Peng
Zhou, Yue
Chaudhari, Shreyas
Kalyan, Ashwin
Rajpurohit, Tanmay
Deshpande, Ameet
Narasimhan, Karthik
Murahari, Vishvak
author_facet Samuel, Vinay
Zou, Henry Peng
Zhou, Yue
Chaudhari, Shreyas
Kalyan, Ashwin
Rajpurohit, Tanmay
Deshpande, Ameet
Narasimhan, Karthik
Murahari, Vishvak
contents Persona agents, which are LLM agents conditioned to act according to an assigned persona, enable contextually rich and user aligned interactions across domains like education and healthcare. However, evaluating how faithfully these agents adhere to their personas remains a significant challenge, particularly in free-form settings that demand consistency across diverse, persona-relevant environments. We introduce PersonaGym, the first dynamic evaluation framework for persona agents, and PersonaScore, a human-aligned automatic metric grounded in decision theory that enables comprehensive large-scale evaluation. Our evaluation of 10 leading LLMs across 200 personas and 10,000 questions reveals significant advancement opportunities. For example, GPT-4.1 had the exact same PersonaScore as LLaMA-3-8b despite being a more recent and advanced closed source model. Importantly, increased model size and complexity do not necessarily enhance persona agent capabilities, underscoring the need for algorithmic and architectural innovation toward faithful, performant persona agents.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18416
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PersonaGym: Evaluating Persona Agents and LLMs
Samuel, Vinay
Zou, Henry Peng
Zhou, Yue
Chaudhari, Shreyas
Kalyan, Ashwin
Rajpurohit, Tanmay
Deshpande, Ameet
Narasimhan, Karthik
Murahari, Vishvak
Computation and Language
Artificial Intelligence
Machine Learning
Persona agents, which are LLM agents conditioned to act according to an assigned persona, enable contextually rich and user aligned interactions across domains like education and healthcare. However, evaluating how faithfully these agents adhere to their personas remains a significant challenge, particularly in free-form settings that demand consistency across diverse, persona-relevant environments. We introduce PersonaGym, the first dynamic evaluation framework for persona agents, and PersonaScore, a human-aligned automatic metric grounded in decision theory that enables comprehensive large-scale evaluation. Our evaluation of 10 leading LLMs across 200 personas and 10,000 questions reveals significant advancement opportunities. For example, GPT-4.1 had the exact same PersonaScore as LLaMA-3-8b despite being a more recent and advanced closed source model. Importantly, increased model size and complexity do not necessarily enhance persona agent capabilities, underscoring the need for algorithmic and architectural innovation toward faithful, performant persona agents.
title PersonaGym: Evaluating Persona Agents and LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.18416