PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Lingfeng, Zhang, Jialing, Gao, Jin, Jiang, Mohan, Wang, Dequan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
di: Jiang, Mohan, et al.
Pubblicazione: (2025)
di: Jiang, Mohan, et al.
Pubblicazione: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
di: Zhou, Yuhang, et al.
Pubblicazione: (2025)
di: Zhou, Yuhang, et al.
Pubblicazione: (2025)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
di: Wang, Xintao, et al.
Pubblicazione: (2025)
di: Wang, Xintao, et al.
Pubblicazione: (2025)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
di: Wang, Kai, et al.
Pubblicazione: (2026)
di: Wang, Kai, et al.
Pubblicazione: (2026)
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
di: Zhang, Qiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Qiyuan, et al.
Pubblicazione: (2024)
Saying the Unsaid: Revealing the Hidden Language of Multimodal Systems Through Telephone Games
di: Zhao, Juntu, et al.
Pubblicazione: (2025)
di: Zhao, Juntu, et al.
Pubblicazione: (2025)
Enhancing Persona Consistency for LLMs' Role-Playing using Persona-Aware Contrastive Learning
di: Ji, Ke, et al.
Pubblicazione: (2025)
di: Ji, Ke, et al.
Pubblicazione: (2025)
Eval4Sim: An Evaluation Framework for Persona Simulation
di: Bao, Eliseo, et al.
Pubblicazione: (2026)
di: Bao, Eliseo, et al.
Pubblicazione: (2026)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
di: Tu, Quan, et al.
Pubblicazione: (2024)
di: Tu, Quan, et al.
Pubblicazione: (2024)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
RepEval: Effective Text Evaluation with LLM Representation
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
di: Yao, Bingsheng, et al.
Pubblicazione: (2025)
di: Yao, Bingsheng, et al.
Pubblicazione: (2025)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
di: Wang, Xiaoyang, et al.
Pubblicazione: (2025)
di: Wang, Xiaoyang, et al.
Pubblicazione: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
di: Lu, Dongxu, et al.
Pubblicazione: (2025)
di: Lu, Dongxu, et al.
Pubblicazione: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
di: de Araujo, Pedro Henrique Luz, et al.
Pubblicazione: (2025)
di: de Araujo, Pedro Henrique Luz, et al.
Pubblicazione: (2025)
Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
di: Tseng, Yu-Min, et al.
Pubblicazione: (2024)
di: Tseng, Yu-Min, et al.
Pubblicazione: (2024)
From Role-Play to Drama-Interaction: An LLM Solution
di: Wu, Weiqi, et al.
Pubblicazione: (2024)
di: Wu, Weiqi, et al.
Pubblicazione: (2024)
From Persona to Personalization: A Survey on Role-Playing Language Agents
di: Chen, Jiangjie, et al.
Pubblicazione: (2024)
di: Chen, Jiangjie, et al.
Pubblicazione: (2024)
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
di: Zhang, Qiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Qiyuan, et al.
Pubblicazione: (2025)
From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents
di: Zhang, Rongsheng, et al.
Pubblicazione: (2026)
di: Zhang, Rongsheng, et al.
Pubblicazione: (2026)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
di: Zhou, Yixi, et al.
Pubblicazione: (2026)
di: Zhou, Yixi, et al.
Pubblicazione: (2026)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
di: D'Souza, Jennifer, et al.
Pubblicazione: (2025)
di: D'Souza, Jennifer, et al.
Pubblicazione: (2025)
SocialBench: Sociality Evaluation of Role-Playing Conversational Agents
di: Chen, Hongzhan, et al.
Pubblicazione: (2024)
di: Chen, Hongzhan, et al.
Pubblicazione: (2024)
Bullying the Machine: How Personas Increase LLM Vulnerability
di: Xu, Ziwei, et al.
Pubblicazione: (2025)
di: Xu, Ziwei, et al.
Pubblicazione: (2025)
BatchEval: Towards Human-like Text Evaluation
di: Yuan, Peiwen, et al.
Pubblicazione: (2023)
di: Yuan, Peiwen, et al.
Pubblicazione: (2023)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
di: Shen, Chengyu, et al.
Pubblicazione: (2026)
di: Shen, Chengyu, et al.
Pubblicazione: (2026)
CharacterGPT: A Persona Reconstruction Framework for Role-Playing Agents
di: Park, Jeiyoon, et al.
Pubblicazione: (2024)
di: Park, Jeiyoon, et al.
Pubblicazione: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
di: Zhou, Xin, et al.
Pubblicazione: (2025)
di: Zhou, Xin, et al.
Pubblicazione: (2025)
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework
di: Yang, Bohao, et al.
Pubblicazione: (2024)
di: Yang, Bohao, et al.
Pubblicazione: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
di: Chen, Junjie, et al.
Pubblicazione: (2026)
di: Chen, Junjie, et al.
Pubblicazione: (2026)
BotEval: Facilitating Interactive Human Evaluation
di: Cho, Hyundong, et al.
Pubblicazione: (2024)
di: Cho, Hyundong, et al.
Pubblicazione: (2024)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
di: Costa, Davi Bastos, et al.
Pubblicazione: (2025)
di: Costa, Davi Bastos, et al.
Pubblicazione: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
di: Han, Steve, et al.
Pubblicazione: (2025)
di: Han, Steve, et al.
Pubblicazione: (2025)
Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs
di: Tang, Wenqiu, et al.
Pubblicazione: (2026)
di: Tang, Wenqiu, et al.
Pubblicazione: (2026)
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
di: Liu, Yuxin, et al.
Pubblicazione: (2026)
di: Liu, Yuxin, et al.
Pubblicazione: (2026)
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
di: Lee, Yukyung, et al.
Pubblicazione: (2024)
di: Lee, Yukyung, et al.
Pubblicazione: (2024)
CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
di: Li, Haitao, et al.
Pubblicazione: (2024)
di: Li, Haitao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
di: Jiang, Mohan, et al.
Pubblicazione: (2025) -
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
di: Zhou, Yuhang, et al.
Pubblicazione: (2025) -
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
di: Wang, Xintao, et al.
Pubblicazione: (2025) -
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
di: Wang, Kai, et al.
Pubblicazione: (2026) -
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
di: Zhang, Qiyuan, et al.
Pubblicazione: (2024)