IMPersona: Evaluating Individual Level LM Impersonation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shi, Quan, Jimenez, Carlos E., Dong, Stephen, Seo, Brian, Yao, Caden, Kelch, Adam, Narasimhan, Karthik |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
par: Shi, Quan, et autres
Publié: (2025)
par: Shi, Quan, et autres
Publié: (2025)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
par: Yang, John, et autres
Publié: (2024)
par: Yang, John, et autres
Publié: (2024)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
par: Kim, Tae Soo, et autres
Publié: (2023)
par: Kim, Tae Soo, et autres
Publié: (2023)
Detection and Positive Reconstruction of Cognitive Distortion sentences: Mandarin Dataset and Evaluation
par: Lin, Shuya, et autres
Publié: (2024)
par: Lin, Shuya, et autres
Publié: (2024)
One Agent Too Many: User Perspectives on Approaches to Multi-agent Conversational AI
par: Clarke, Christopher, et autres
Publié: (2024)
par: Clarke, Christopher, et autres
Publié: (2024)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
par: Zhou, Kaitlyn, et autres
Publié: (2024)
par: Zhou, Kaitlyn, et autres
Publié: (2024)
Using a Human-AI Teaming Approach to Create and Curate Scientific Datasets with the SCILIRE System
par: Bölücü, Necva, et autres
Publié: (2026)
par: Bölücü, Necva, et autres
Publié: (2026)
CloChat: Understanding How People Customize, Interact, and Experience Personas in Large Language Models
par: Ha, Juhye, et autres
Publié: (2024)
par: Ha, Juhye, et autres
Publié: (2024)
Generating Educational Materials with Different Levels of Readability using LLMs
par: Huang, Chieh-Yang, et autres
Publié: (2024)
par: Huang, Chieh-Yang, et autres
Publié: (2024)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
par: Liu, Tianjian, et autres
Publié: (2025)
par: Liu, Tianjian, et autres
Publié: (2025)
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
par: Chen, Chaoran, et autres
Publié: (2025)
par: Chen, Chaoran, et autres
Publié: (2025)
BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
par: Yu, Yeyong, et autres
Publié: (2024)
par: Yu, Yeyong, et autres
Publié: (2024)
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models
par: Wang, Zhiyuan, et autres
Publié: (2024)
par: Wang, Zhiyuan, et autres
Publié: (2024)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
par: Sun, Lu, et autres
Publié: (2025)
par: Sun, Lu, et autres
Publié: (2025)
Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models
par: Dhillon, Paramveer S., et autres
Publié: (2024)
par: Dhillon, Paramveer S., et autres
Publié: (2024)
Designing KRIYA: An AI Companion for Wellbeing Self-Reflection
par: Zhu, Shanshan, et autres
Publié: (2026)
par: Zhu, Shanshan, et autres
Publié: (2026)
Disentangling Prompt Element Level Risk Factors for Hallucinations and Omissions in Mental Health LLM Responses
par: Ni, Congning, et autres
Publié: (2026)
par: Ni, Congning, et autres
Publié: (2026)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
par: Vo, Truong, et autres
Publié: (2025)
par: Vo, Truong, et autres
Publié: (2025)
KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
par: Nayeem, Mir Tafseer, et autres
Publié: (2024)
par: Nayeem, Mir Tafseer, et autres
Publié: (2024)
Investigating User Perspectives on Differentially Private Text Privatization
par: Meisenbacher, Stephen, et autres
Publié: (2025)
par: Meisenbacher, Stephen, et autres
Publié: (2025)
Navigating the Landscape of Hint Generation Research: From the Past to the Future
par: Jangra, Anubhav, et autres
Publié: (2024)
par: Jangra, Anubhav, et autres
Publié: (2024)
Robots in the Middle: Evaluating LLMs in Dispute Resolution
par: Tan, Jinzhe, et autres
Publié: (2024)
par: Tan, Jinzhe, et autres
Publié: (2024)
Pearmut: Human Evaluation of Translation Made Trivial
par: Zouhar, Vilém, et autres
Publié: (2026)
par: Zouhar, Vilém, et autres
Publié: (2026)
Collage: Decomposable Rapid Prototyping for Information Extraction on Scientific PDFs
par: Gururaja, Sireesh, et autres
Publié: (2024)
par: Gururaja, Sireesh, et autres
Publié: (2024)
WordCraft: Scaffolding the Keyword Method for L2 Vocabulary Learning with Multimodal LLMs
par: Shao, Yuheng, et autres
Publié: (2026)
par: Shao, Yuheng, et autres
Publié: (2026)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
par: Shin, Jisu, et autres
Publié: (2025)
par: Shin, Jisu, et autres
Publié: (2025)
Explainable AI Components for Narrative Map Extraction
par: Keith, Brian, et autres
Publié: (2025)
par: Keith, Brian, et autres
Publié: (2025)
Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models
par: Tran, Son Quoc, et autres
Publié: (2025)
par: Tran, Son Quoc, et autres
Publié: (2025)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
par: Liang, Chen, et autres
Publié: (2026)
par: Liang, Chen, et autres
Publié: (2026)
Context-Aware Monolingual Human Evaluation of Machine Translation
par: Picinini, Silvio, et autres
Publié: (2025)
par: Picinini, Silvio, et autres
Publié: (2025)
Designing and Evaluating Chain-of-Hints for Scientific Question Answering
par: Jangra, Anubhav, et autres
Publié: (2025)
par: Jangra, Anubhav, et autres
Publié: (2025)
Aligning LLMs with Individual Preferences via Interaction
par: Wu, Shujin, et autres
Publié: (2024)
par: Wu, Shujin, et autres
Publié: (2024)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
par: Shaikh, Omar, et autres
Publié: (2025)
par: Shaikh, Omar, et autres
Publié: (2025)
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
par: Joshi, Brihi, et autres
Publié: (2025)
par: Joshi, Brihi, et autres
Publié: (2025)
DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
par: Shrivastava, Aryan, et autres
Publié: (2025)
par: Shrivastava, Aryan, et autres
Publié: (2025)
Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality
par: Haq, Sami Ul, et autres
Publié: (2025)
par: Haq, Sami Ul, et autres
Publié: (2025)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
par: Lee, Suhyun, et autres
Publié: (2026)
par: Lee, Suhyun, et autres
Publié: (2026)
Human-LLM Collaborative Construction of a Cantonese Emotion Lexicon
par: Zhang, Yusong, et autres
Publié: (2024)
par: Zhang, Yusong, et autres
Publié: (2024)
Evaluation of a Sign Language Avatar on Comprehensibility, User Experience \& Acceptability
par: Wasserroth, Fenya, et autres
Publié: (2025)
par: Wasserroth, Fenya, et autres
Publié: (2025)
IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data
par: Peng, Bo, et autres
Publié: (2025)
par: Peng, Bo, et autres
Publié: (2025)
Documents similaires
-
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
par: Shi, Quan, et autres
Publié: (2025) -
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
par: Yang, John, et autres
Publié: (2024) -
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
par: Kim, Tae Soo, et autres
Publié: (2023) -
Detection and Positive Reconstruction of Cognitive Distortion sentences: Mandarin Dataset and Evaluation
par: Lin, Shuya, et autres
Publié: (2024) -
One Agent Too Many: User Perspectives on Approaches to Multi-agent Conversational AI
par: Clarke, Christopher, et autres
Publié: (2024)