PersonaGym: Evaluating Persona Agents and LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Samuel, Vinay, Zou, Henry Peng, Zhou, Yue, Chaudhari, Shreyas, Kalyan, Ashwin, Rajpurohit, Tanmay, Deshpande, Ameet, Narasimhan, Karthik, Murahari, Vishvak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing AI Safety with Source Code
by: Narayan, Ujwal, et al.
Published: (2025)
by: Narayan, Ujwal, et al.
Published: (2025)
GEO: Generative Engine Optimization
by: Aggarwal, Pranjal, et al.
Published: (2023)
by: Aggarwal, Pranjal, et al.
Published: (2023)
Agent Context Protocols Enhance Collective Inference
by: Bhardwaj, Devansh, et al.
Published: (2025)
by: Bhardwaj, Devansh, et al.
Published: (2025)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
QualEval: Qualitative Evaluation for Model Improvement
by: Murahari, Vishvak, et al.
Published: (2023)
by: Murahari, Vishvak, et al.
Published: (2023)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
by: Gupta, Shashank, et al.
Published: (2023)
by: Gupta, Shashank, et al.
Published: (2023)
Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation
by: Dogra, Atharvan, et al.
Published: (2024)
by: Dogra, Atharvan, et al.
Published: (2024)
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
by: Dogra, Atharvan, et al.
Published: (2025)
by: Dogra, Atharvan, et al.
Published: (2025)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
by: Li, Wenkai, et al.
Published: (2026)
by: Li, Wenkai, et al.
Published: (2026)
PersonaBOT: Bringing Customer Personas to Life with LLMs and RAG
by: Rizwan, Muhammed, et al.
Published: (2025)
by: Rizwan, Muhammed, et al.
Published: (2025)
From Guessing to Asking: An Approach to Resolving the Persona Knowledge Gap in LLMs during Multi-Turn Conversations
by: Baskar, Sarvesh, et al.
Published: (2025)
by: Baskar, Sarvesh, et al.
Published: (2025)
Dimension-Free Parameterized Approximation Schemes for Hybrid Clustering
by: Gadekar, Ameet, et al.
Published: (2025)
by: Gadekar, Ameet, et al.
Published: (2025)
LLMs + Persona-Plug = Personalized LLMs
by: Liu, Jiongnan, et al.
Published: (2024)
by: Liu, Jiongnan, et al.
Published: (2024)
Localizing Persona Representations in LLMs
by: Cintas, Celia, et al.
Published: (2025)
by: Cintas, Celia, et al.
Published: (2025)
PersonaMatrix: A Recipe for Persona-Aware Evaluation of Legal Summarization
by: Pang, Tsz Fung, et al.
Published: (2025)
by: Pang, Tsz Fung, et al.
Published: (2025)
Enhancing Persona Consistency for LLMs' Role-Playing using Persona-Aware Contrastive Learning
by: Ji, Ke, et al.
Published: (2025)
by: Ji, Ke, et al.
Published: (2025)
LLMs' ways of seeing User Personas
by: Panda, Swaroop
Published: (2024)
by: Panda, Swaroop
Published: (2024)
Styles + Persona-plug = Customized LLMs
by: Song, Yutong, et al.
Published: (2026)
by: Song, Yutong, et al.
Published: (2026)
PersonaLedger: Generating Realistic Financial Transactions with Persona Conditioned LLMs and Rule Grounded Feedback
by: Yuan, Dehao, et al.
Published: (2026)
by: Yuan, Dehao, et al.
Published: (2026)
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models
by: Shi, Wenlong, et al.
Published: (2026)
by: Shi, Wenlong, et al.
Published: (2026)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
by: Truong, Kimberly Le, et al.
Published: (2025)
by: Truong, Kimberly Le, et al.
Published: (2025)
Simulating Misinformation Vulnerabilities With Agent Personas
by: Farr, David, et al.
Published: (2025)
by: Farr, David, et al.
Published: (2025)
Persona
Published: (2024)
Published: (2024)
NARRA-Gym for Evaluating Interactive Narrative Agents
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025)
by: Wang, Zhen, et al.
Published: (2025)
ImplicitAVE: An Open-Source Dataset and Multimodal LLMs Benchmark for Implicit Attribute Value Extraction
by: Zou, Henry Peng, et al.
Published: (2024)
by: Zou, Henry Peng, et al.
Published: (2024)
PersonaAgent with GraphRAG: Community-Aware Knowledge Graphs for Personalized LLM
by: Liang, Siqi, et al.
Published: (2025)
by: Liang, Siqi, et al.
Published: (2025)
SimPersona: Learning Discrete Buyer Personas from Raw Clickstreams for Grounded E-Commerce Agents
by: Foumani, Zahra Zanjani, et al.
Published: (2026)
by: Foumani, Zahra Zanjani, et al.
Published: (2026)
Clustering under Constraints: Efficient Parameterized Approximation Schemes
by: Bhore, Sujoy, et al.
Published: (2025)
by: Bhore, Sujoy, et al.
Published: (2025)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
Evaluation of LLMs Biases Towards Elite Universities: A Persona-Based Exploration
by: Gupta, Shailja, et al.
Published: (2024)
by: Gupta, Shailja, et al.
Published: (2024)
Persona-based Multi-Agent Collaboration for Brainstorming
by: Straub, Nate, et al.
Published: (2025)
by: Straub, Nate, et al.
Published: (2025)
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework
by: Yang, Bohao, et al.
Published: (2024)
by: Yang, Bohao, et al.
Published: (2024)
Similar Items
-
Probing AI Safety with Source Code
by: Narayan, Ujwal, et al.
Published: (2025) -
GEO: Generative Engine Optimization
by: Aggarwal, Pranjal, et al.
Published: (2023) -
Agent Context Protocols Enhance Collective Inference
by: Bhardwaj, Devansh, et al.
Published: (2025) -
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
by: Chaudhari, Shreyas, et al.
Published: (2024) -
QualEval: Qualitative Evaluation for Model Improvement
by: Murahari, Vishvak, et al.
Published: (2023)