Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Dongxu, Jeuring, Johan, Gatt, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
by: Li, Nan, et al.
Published: (2025)
by: Li, Nan, et al.
Published: (2025)
Contrast Is All You Need
by: Kilic, Burak, et al.
Published: (2023)
by: Kilic, Burak, et al.
Published: (2023)
How and where does CLIP process negation?
by: Quantmeyer, Vincent, et al.
Published: (2024)
by: Quantmeyer, Vincent, et al.
Published: (2024)
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
by: Zhu, Ruizhe, et al.
Published: (2025)
by: Zhu, Ruizhe, et al.
Published: (2025)
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
by: Ou, Jiao, et al.
Published: (2023)
by: Ou, Jiao, et al.
Published: (2023)
DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
by: Yao, Bingsheng, et al.
Published: (2025)
by: Yao, Bingsheng, et al.
Published: (2025)
Role-Playing Evaluation for Large Language Models
by: Boudouri, Yassine El, et al.
Published: (2025)
by: Boudouri, Yassine El, et al.
Published: (2025)
Collaborative Storytelling and LLM: A Linguistic Analysis of Automatically-Generated Role-Playing Game Sessions
by: Maisto, Alessandro
Published: (2025)
by: Maisto, Alessandro
Published: (2025)
Improving LLM Reasoning through Interpretable Role-Playing Steering
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained Counselors
by: Qi, Zhiyang, et al.
Published: (2025)
by: Qi, Zhiyang, et al.
Published: (2025)
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
by: Wang, Zekun Moore, et al.
Published: (2023)
by: Wang, Zekun Moore, et al.
Published: (2023)
FedDTRE: Federated Dialogue Generation Models Powered by Trustworthiness Evaluation
by: Lu, Shule, et al.
Published: (2025)
by: Lu, Shule, et al.
Published: (2025)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
by: Wang, Xintao, et al.
Published: (2025)
by: Wang, Xintao, et al.
Published: (2025)
LLM Discussion: Enhancing the Creativity of Large Language Models via Discussion Framework and Role-Play
by: Lu, Li-Chun, et al.
Published: (2024)
by: Lu, Li-Chun, et al.
Published: (2024)
Response Enhanced Semi-supervised Dialogue Query Generation
by: Huang, Jianheng, et al.
Published: (2023)
by: Huang, Jianheng, et al.
Published: (2023)
Enhancing Hallucination Detection through Perturbation-Based Synthetic Data Generation in System Responses
by: Zhang, Dongxu, et al.
Published: (2024)
by: Zhang, Dongxu, et al.
Published: (2024)
Emotional Support with LLM-based Empathetic Dialogue Generation
by: Wang, Shiquan, et al.
Published: (2025)
by: Wang, Shiquan, et al.
Published: (2025)
Synthetic Dialogue Dataset Generation using LLM Agents
by: Abdullin, Yelaman, et al.
Published: (2024)
by: Abdullin, Yelaman, et al.
Published: (2024)
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
by: Yu, Pengfei, et al.
Published: (2025)
by: Yu, Pengfei, et al.
Published: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
by: Yao, Jing, et al.
Published: (2024)
by: Yao, Jing, et al.
Published: (2024)
Plug-and-Play Policy Planner for Large Language Model Powered Dialogue Agents
by: Deng, Yang, et al.
Published: (2023)
by: Deng, Yang, et al.
Published: (2023)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
by: Wu, Weihao, et al.
Published: (2025)
by: Wu, Weihao, et al.
Published: (2025)
Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
Dialogue You Can Trust: Human and AI Perspectives on Generated Conversations
by: Ebubechukwu, Ike, et al.
Published: (2024)
by: Ebubechukwu, Ike, et al.
Published: (2024)
PersonaKit (PK): A Plug-and-Play Platform for User Testing Diverse Roles in Full-Duplex Dialogue
by: Jeon, Hyunbae, et al.
Published: (2026)
by: Jeon, Hyunbae, et al.
Published: (2026)
Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules
by: Zhang, Yueqi, et al.
Published: (2025)
by: Zhang, Yueqi, et al.
Published: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents
by: Zhang, Rongsheng, et al.
Published: (2026)
by: Zhang, Rongsheng, et al.
Published: (2026)
Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects
by: Peng, Ji-Lun, et al.
Published: (2026)
by: Peng, Ji-Lun, et al.
Published: (2026)
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses
by: Lu, Xiaotian, et al.
Published: (2024)
by: Lu, Xiaotian, et al.
Published: (2024)
Mind the Gap: The Divergence Between Human and LLM-Generated Tasks
by: Lu, Yi-Long, et al.
Published: (2025)
by: Lu, Yi-Long, et al.
Published: (2025)
RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents
by: Rosati, Riccardo, et al.
Published: (2026)
by: Rosati, Riccardo, et al.
Published: (2026)
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
by: Choi, Junhyuk, et al.
Published: (2026)
by: Choi, Junhyuk, et al.
Published: (2026)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing
by: Xu, Zhenhua, et al.
Published: (2026)
by: Xu, Zhenhua, et al.
Published: (2026)
Developing A Framework to Support Human Evaluation of Bias in Generated Free Response Text
by: Healey, Jennifer, et al.
Published: (2025)
by: Healey, Jennifer, et al.
Published: (2025)
Decision-Oriented Dialogue for Human-AI Collaboration
by: Lin, Jessy, et al.
Published: (2023)
by: Lin, Jessy, et al.
Published: (2023)
MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues
by: Binici, Kuluhan, et al.
Published: (2024)
by: Binici, Kuluhan, et al.
Published: (2024)
Similar Items
-
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
by: Li, Nan, et al.
Published: (2025) -
Contrast Is All You Need
by: Kilic, Burak, et al.
Published: (2023) -
How and where does CLIP process negation?
by: Quantmeyer, Vincent, et al.
Published: (2024) -
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
by: Zhu, Ruizhe, et al.
Published: (2025) -
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
by: Ou, Jiao, et al.
Published: (2023)