Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Shin, Jisu, Oh, Juhyun, Kim, Eunsu, Song, Hoyun, Oh, Alice |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
por: Oh, Juhyun, et al.
Publicado: (2024)
por: Oh, Juhyun, et al.
Publicado: (2024)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
por: Shin, Jisu, et al.
Publicado: (2025)
por: Shin, Jisu, et al.
Publicado: (2025)
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
por: Oh, Juhyun, et al.
Publicado: (2025)
por: Oh, Juhyun, et al.
Publicado: (2025)
CloChat: Understanding How People Customize, Interact, and Experience Personas in Large Language Models
por: Ha, Juhye, et al.
Publicado: (2024)
por: Ha, Juhye, et al.
Publicado: (2024)
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
por: Shafayat, Sheikh, et al.
Publicado: (2024)
por: Shafayat, Sheikh, et al.
Publicado: (2024)
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
por: Wang, Ziyi, et al.
Publicado: (2025)
por: Wang, Ziyi, et al.
Publicado: (2025)
Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
por: Amin, Danial, et al.
Publicado: (2025)
por: Amin, Danial, et al.
Publicado: (2025)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
por: Yoon, Sion, et al.
Publicado: (2024)
por: Yoon, Sion, et al.
Publicado: (2024)
One-Topic-Doesn't-Fit-All: Transcreating Reading Comprehension Test for Personalized Learning
por: Han, Jieun, et al.
Publicado: (2025)
por: Han, Jieun, et al.
Publicado: (2025)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
por: Afzoon, Saleh, et al.
Publicado: (2026)
por: Afzoon, Saleh, et al.
Publicado: (2026)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
por: Afzoon, Saleh, et al.
Publicado: (2026)
por: Afzoon, Saleh, et al.
Publicado: (2026)
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
por: Lee, Yoonjoo, et al.
Publicado: (2024)
por: Lee, Yoonjoo, et al.
Publicado: (2024)
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
por: Fröhling, Leon, et al.
Publicado: (2024)
por: Fröhling, Leon, et al.
Publicado: (2024)
Dialogue Language Model with Large-Scale Persona Data Engineering
por: Hong, Mengze, et al.
Publicado: (2024)
por: Hong, Mengze, et al.
Publicado: (2024)
Blind Spots and Biases: Exploring the Role of Annotator Cognitive Biases in NLP
por: Gautam, Sanjana, et al.
Publicado: (2024)
por: Gautam, Sanjana, et al.
Publicado: (2024)
The Hall of Singularity: VR Experience of Prophecy by AI
por: Kim, Jisu, et al.
Publicado: (2024)
por: Kim, Jisu, et al.
Publicado: (2024)
Adapting Large Language Models for Character-based Augmentative and Alternative Communication
por: Gaines, Dylan, et al.
Publicado: (2025)
por: Gaines, Dylan, et al.
Publicado: (2025)
ReDemon UI: Reactive Synthesis by Demonstration for Web UI
por: Lee, Jay, et al.
Publicado: (2025)
por: Lee, Jay, et al.
Publicado: (2025)
Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-Playing
por: Zhang, Wenyuan, et al.
Publicado: (2024)
por: Zhang, Wenyuan, et al.
Publicado: (2024)
MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education
por: Yue, Murong, et al.
Publicado: (2024)
por: Yue, Murong, et al.
Publicado: (2024)
Predicting and Understanding Turn-Taking Behavior in Open-Ended Group Activities in Virtual Reality
por: Wang, Portia, et al.
Publicado: (2024)
por: Wang, Portia, et al.
Publicado: (2024)
VicSim: Enhancing Victim Simulation with Emotional and Linguistic Fidelity
por: Li, Yerong, et al.
Publicado: (2025)
por: Li, Yerong, et al.
Publicado: (2025)
Generating Educational Materials with Different Levels of Readability using LLMs
por: Huang, Chieh-Yang, et al.
Publicado: (2024)
por: Huang, Chieh-Yang, et al.
Publicado: (2024)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
por: Kim, Tae Soo, et al.
Publicado: (2023)
por: Kim, Tae Soo, et al.
Publicado: (2023)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
por: Oh, Juhyun, et al.
Publicado: (2024)
por: Oh, Juhyun, et al.
Publicado: (2024)
The Adoption and Efficacy of Large Language Models: Evidence From Consumer Complaints in the Financial Industry
por: Shin, Minkyu, et al.
Publicado: (2023)
por: Shin, Minkyu, et al.
Publicado: (2023)
Does Difficulty even Matter? Investigating Difficulty Adjustment and Practice Behavior in an Open-Ended Learning Task
por: Schütt, Anan, et al.
Publicado: (2023)
por: Schütt, Anan, et al.
Publicado: (2023)
How Neurotypical and Autistic Children Interact Nonverbally with Anthropomorphic Agents in Open-Ended Tasks
por: Zhang, Chuxuan, et al.
Publicado: (2026)
por: Zhang, Chuxuan, et al.
Publicado: (2026)
Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task Automation
por: Lee, Sunjae, et al.
Publicado: (2023)
por: Lee, Sunjae, et al.
Publicado: (2023)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
por: Lee, Jungjae, et al.
Publicado: (2025)
por: Lee, Jungjae, et al.
Publicado: (2025)
SensorPersona: An LLM-Empowered System for Continual Persona Extraction from Longitudinal Mobile Sensor Streams
por: Yang, Bufang, et al.
Publicado: (2026)
por: Yang, Bufang, et al.
Publicado: (2026)
Explainable XR: Understanding User Behaviors of XR Environments using LLM-assisted Analytics Framework
por: Kim, Yoonsang, et al.
Publicado: (2025)
por: Kim, Yoonsang, et al.
Publicado: (2025)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
por: Kim, Eunsu, et al.
Publicado: (2025)
por: Kim, Eunsu, et al.
Publicado: (2025)
CHOP: Integrating ChatGPT into EFL Oral Presentation Practice
por: Cha, Jungyoub, et al.
Publicado: (2024)
por: Cha, Jungyoub, et al.
Publicado: (2024)
I am not thinking anymore, just following the path.: Investigating Task Delegation Trend of Author-AI Co-Creation with Generative AIs
por: Kim, Yujin, et al.
Publicado: (2025)
por: Kim, Yujin, et al.
Publicado: (2025)
When Scaffolding Breaks: Investigating Student Interaction with LLM-Based Writing Support in Real-Time K-12 EFL Classrooms
por: Myung, Junho, et al.
Publicado: (2025)
por: Myung, Junho, et al.
Publicado: (2025)
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
por: Kim, Taesoo, et al.
Publicado: (2025)
por: Kim, Taesoo, et al.
Publicado: (2025)
Supporting the Digital Autonomy of Elders Through LLM Assistance
por: Roberts, Jesse, et al.
Publicado: (2024)
por: Roberts, Jesse, et al.
Publicado: (2024)
Cheap and Easy Open-Ended Text Input for Interactive Emergent Narrative
por: Kreminski, Max
Publicado: (2024)
por: Kreminski, Max
Publicado: (2024)
Evaluating LLM-Generated Q&A Test: a Student-Centered Study
por: Wróblewska, Anna, et al.
Publicado: (2025)
por: Wróblewska, Anna, et al.
Publicado: (2025)
Ejemplares similares
-
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
por: Oh, Juhyun, et al.
Publicado: (2024) -
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
por: Shin, Jisu, et al.
Publicado: (2025) -
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
por: Oh, Juhyun, et al.
Publicado: (2025) -
CloChat: Understanding How People Customize, Interact, and Experience Personas in Large Language Models
por: Ha, Juhye, et al.
Publicado: (2024) -
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
por: Shafayat, Sheikh, et al.
Publicado: (2024)