PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhou, Lingfeng, Zhang, Jialing, Gao, Jin, Jiang, Mohan, Wang, Dequan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
par: Jiang, Mohan, et autres
Publié: (2025)
par: Jiang, Mohan, et autres
Publié: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
par: Zhou, Yuhang, et autres
Publié: (2025)
par: Zhou, Yuhang, et autres
Publié: (2025)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
par: Wang, Xintao, et autres
Publié: (2025)
par: Wang, Xintao, et autres
Publié: (2025)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
par: Wang, Kai, et autres
Publié: (2026)
par: Wang, Kai, et autres
Publié: (2026)
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
par: Zhang, Qiyuan, et autres
Publié: (2024)
par: Zhang, Qiyuan, et autres
Publié: (2024)
Saying the Unsaid: Revealing the Hidden Language of Multimodal Systems Through Telephone Games
par: Zhao, Juntu, et autres
Publié: (2025)
par: Zhao, Juntu, et autres
Publié: (2025)
Enhancing Persona Consistency for LLMs' Role-Playing using Persona-Aware Contrastive Learning
par: Ji, Ke, et autres
Publié: (2025)
par: Ji, Ke, et autres
Publié: (2025)
Eval4Sim: An Evaluation Framework for Persona Simulation
par: Bao, Eliseo, et autres
Publié: (2026)
par: Bao, Eliseo, et autres
Publié: (2026)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
par: Tu, Quan, et autres
Publié: (2024)
par: Tu, Quan, et autres
Publié: (2024)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
par: Son, Guijin, et autres
Publié: (2024)
par: Son, Guijin, et autres
Publié: (2024)
RepEval: Effective Text Evaluation with LLM Representation
par: Sheng, Shuqian, et autres
Publié: (2024)
par: Sheng, Shuqian, et autres
Publié: (2024)
DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
par: Yao, Bingsheng, et autres
Publié: (2025)
par: Yao, Bingsheng, et autres
Publié: (2025)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
par: Wang, Xiaoyang, et autres
Publié: (2025)
par: Wang, Xiaoyang, et autres
Publié: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
par: Yang, Langqi, et autres
Publié: (2025)
par: Yang, Langqi, et autres
Publié: (2025)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
par: Lu, Dongxu, et autres
Publié: (2025)
par: Lu, Dongxu, et autres
Publié: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
par: Zhou, Yilun, et autres
Publié: (2025)
par: Zhou, Yilun, et autres
Publié: (2025)
Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
par: de Araujo, Pedro Henrique Luz, et autres
Publié: (2025)
par: de Araujo, Pedro Henrique Luz, et autres
Publié: (2025)
Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
par: Tseng, Yu-Min, et autres
Publié: (2024)
par: Tseng, Yu-Min, et autres
Publié: (2024)
From Role-Play to Drama-Interaction: An LLM Solution
par: Wu, Weiqi, et autres
Publié: (2024)
par: Wu, Weiqi, et autres
Publié: (2024)
From Persona to Personalization: A Survey on Role-Playing Language Agents
par: Chen, Jiangjie, et autres
Publié: (2024)
par: Chen, Jiangjie, et autres
Publié: (2024)
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
par: Zhang, Qiyuan, et autres
Publié: (2025)
par: Zhang, Qiyuan, et autres
Publié: (2025)
From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents
par: Zhang, Rongsheng, et autres
Publié: (2026)
par: Zhang, Rongsheng, et autres
Publié: (2026)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
par: Zhou, Yixi, et autres
Publié: (2026)
par: Zhou, Yixi, et autres
Publié: (2026)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
par: Mohammadi, Hadi, et autres
Publié: (2025)
par: Mohammadi, Hadi, et autres
Publié: (2025)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
par: D'Souza, Jennifer, et autres
Publié: (2025)
par: D'Souza, Jennifer, et autres
Publié: (2025)
SocialBench: Sociality Evaluation of Role-Playing Conversational Agents
par: Chen, Hongzhan, et autres
Publié: (2024)
par: Chen, Hongzhan, et autres
Publié: (2024)
Bullying the Machine: How Personas Increase LLM Vulnerability
par: Xu, Ziwei, et autres
Publié: (2025)
par: Xu, Ziwei, et autres
Publié: (2025)
BatchEval: Towards Human-like Text Evaluation
par: Yuan, Peiwen, et autres
Publié: (2023)
par: Yuan, Peiwen, et autres
Publié: (2023)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
par: Shen, Chengyu, et autres
Publié: (2026)
par: Shen, Chengyu, et autres
Publié: (2026)
CharacterGPT: A Persona Reconstruction Framework for Role-Playing Agents
par: Park, Jeiyoon, et autres
Publié: (2024)
par: Park, Jeiyoon, et autres
Publié: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
par: Zhou, Xin, et autres
Publié: (2025)
par: Zhou, Xin, et autres
Publié: (2025)
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework
par: Yang, Bohao, et autres
Publié: (2024)
par: Yang, Bohao, et autres
Publié: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
par: Chen, Junjie, et autres
Publié: (2026)
par: Chen, Junjie, et autres
Publié: (2026)
BotEval: Facilitating Interactive Human Evaluation
par: Cho, Hyundong, et autres
Publié: (2024)
par: Cho, Hyundong, et autres
Publié: (2024)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
par: Costa, Davi Bastos, et autres
Publié: (2025)
par: Costa, Davi Bastos, et autres
Publié: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
par: Han, Steve, et autres
Publié: (2025)
par: Han, Steve, et autres
Publié: (2025)
Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs
par: Tang, Wenqiu, et autres
Publié: (2026)
par: Tang, Wenqiu, et autres
Publié: (2026)
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
par: Liu, Yuxin, et autres
Publié: (2026)
par: Liu, Yuxin, et autres
Publié: (2026)
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
par: Lee, Yukyung, et autres
Publié: (2024)
par: Lee, Yukyung, et autres
Publié: (2024)
CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
par: Li, Haitao, et autres
Publié: (2024)
par: Li, Haitao, et autres
Publié: (2024)
Documents similaires
-
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
par: Jiang, Mohan, et autres
Publié: (2025) -
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
par: Zhou, Yuhang, et autres
Publié: (2025) -
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
par: Wang, Xintao, et autres
Publié: (2025) -
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
par: Wang, Kai, et autres
Publié: (2026) -
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
par: Zhang, Qiyuan, et autres
Publié: (2024)