IMPersona: Evaluating Individual Level LM Impersonation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Quan, Jimenez, Carlos E., Dong, Stephen, Seo, Brian, Yao, Caden, Kelch, Adam, Narasimhan, Karthik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
von: Shi, Quan, et al.
Veröffentlicht: (2025)
von: Shi, Quan, et al.
Veröffentlicht: (2025)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
Detection and Positive Reconstruction of Cognitive Distortion sentences: Mandarin Dataset and Evaluation
von: Lin, Shuya, et al.
Veröffentlicht: (2024)
von: Lin, Shuya, et al.
Veröffentlicht: (2024)
One Agent Too Many: User Perspectives on Approaches to Multi-agent Conversational AI
von: Clarke, Christopher, et al.
Veröffentlicht: (2024)
von: Clarke, Christopher, et al.
Veröffentlicht: (2024)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2024)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2024)
Using a Human-AI Teaming Approach to Create and Curate Scientific Datasets with the SCILIRE System
von: Bölücü, Necva, et al.
Veröffentlicht: (2026)
von: Bölücü, Necva, et al.
Veröffentlicht: (2026)
CloChat: Understanding How People Customize, Interact, and Experience Personas in Large Language Models
von: Ha, Juhye, et al.
Veröffentlicht: (2024)
von: Ha, Juhye, et al.
Veröffentlicht: (2024)
Generating Educational Materials with Different Levels of Readability using LLMs
von: Huang, Chieh-Yang, et al.
Veröffentlicht: (2024)
von: Huang, Chieh-Yang, et al.
Veröffentlicht: (2024)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
von: Liu, Tianjian, et al.
Veröffentlicht: (2025)
von: Liu, Tianjian, et al.
Veröffentlicht: (2025)
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
von: Yu, Yeyong, et al.
Veröffentlicht: (2024)
von: Yu, Yeyong, et al.
Veröffentlicht: (2024)
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2024)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
von: Sun, Lu, et al.
Veröffentlicht: (2025)
von: Sun, Lu, et al.
Veröffentlicht: (2025)
Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models
von: Dhillon, Paramveer S., et al.
Veröffentlicht: (2024)
von: Dhillon, Paramveer S., et al.
Veröffentlicht: (2024)
Designing KRIYA: An AI Companion for Wellbeing Self-Reflection
von: Zhu, Shanshan, et al.
Veröffentlicht: (2026)
von: Zhu, Shanshan, et al.
Veröffentlicht: (2026)
Disentangling Prompt Element Level Risk Factors for Hallucinations and Omissions in Mental Health LLM Responses
von: Ni, Congning, et al.
Veröffentlicht: (2026)
von: Ni, Congning, et al.
Veröffentlicht: (2026)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
von: Vo, Truong, et al.
Veröffentlicht: (2025)
von: Vo, Truong, et al.
Veröffentlicht: (2025)
KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2024)
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2024)
Investigating User Perspectives on Differentially Private Text Privatization
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
Navigating the Landscape of Hint Generation Research: From the Past to the Future
von: Jangra, Anubhav, et al.
Veröffentlicht: (2024)
von: Jangra, Anubhav, et al.
Veröffentlicht: (2024)
Robots in the Middle: Evaluating LLMs in Dispute Resolution
von: Tan, Jinzhe, et al.
Veröffentlicht: (2024)
von: Tan, Jinzhe, et al.
Veröffentlicht: (2024)
Pearmut: Human Evaluation of Translation Made Trivial
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
Collage: Decomposable Rapid Prototyping for Information Extraction on Scientific PDFs
von: Gururaja, Sireesh, et al.
Veröffentlicht: (2024)
von: Gururaja, Sireesh, et al.
Veröffentlicht: (2024)
WordCraft: Scaffolding the Keyword Method for L2 Vocabulary Learning with Multimodal LLMs
von: Shao, Yuheng, et al.
Veröffentlicht: (2026)
von: Shao, Yuheng, et al.
Veröffentlicht: (2026)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
Explainable AI Components for Narrative Map Extraction
von: Keith, Brian, et al.
Veröffentlicht: (2025)
von: Keith, Brian, et al.
Veröffentlicht: (2025)
Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models
von: Tran, Son Quoc, et al.
Veröffentlicht: (2025)
von: Tran, Son Quoc, et al.
Veröffentlicht: (2025)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
von: Liang, Chen, et al.
Veröffentlicht: (2026)
von: Liang, Chen, et al.
Veröffentlicht: (2026)
Context-Aware Monolingual Human Evaluation of Machine Translation
von: Picinini, Silvio, et al.
Veröffentlicht: (2025)
von: Picinini, Silvio, et al.
Veröffentlicht: (2025)
Designing and Evaluating Chain-of-Hints for Scientific Question Answering
von: Jangra, Anubhav, et al.
Veröffentlicht: (2025)
von: Jangra, Anubhav, et al.
Veröffentlicht: (2025)
Aligning LLMs with Individual Preferences via Interaction
von: Wu, Shujin, et al.
Veröffentlicht: (2024)
von: Wu, Shujin, et al.
Veröffentlicht: (2024)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
von: Shaikh, Omar, et al.
Veröffentlicht: (2025)
von: Shaikh, Omar, et al.
Veröffentlicht: (2025)
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
von: Joshi, Brihi, et al.
Veröffentlicht: (2025)
von: Joshi, Brihi, et al.
Veröffentlicht: (2025)
DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality
von: Haq, Sami Ul, et al.
Veröffentlicht: (2025)
von: Haq, Sami Ul, et al.
Veröffentlicht: (2025)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
von: Lee, Suhyun, et al.
Veröffentlicht: (2026)
von: Lee, Suhyun, et al.
Veröffentlicht: (2026)
Human-LLM Collaborative Construction of a Cantonese Emotion Lexicon
von: Zhang, Yusong, et al.
Veröffentlicht: (2024)
von: Zhang, Yusong, et al.
Veröffentlicht: (2024)
Evaluation of a Sign Language Avatar on Comprehensibility, User Experience \& Acceptability
von: Wasserroth, Fenya, et al.
Veröffentlicht: (2025)
von: Wasserroth, Fenya, et al.
Veröffentlicht: (2025)
IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data
von: Peng, Bo, et al.
Veröffentlicht: (2025)
von: Peng, Bo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
von: Shi, Quan, et al.
Veröffentlicht: (2025) -
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2024) -
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023) -
Detection and Positive Reconstruction of Cognitive Distortion sentences: Mandarin Dataset and Evaluation
von: Lin, Shuya, et al.
Veröffentlicht: (2024) -
One Agent Too Many: User Perspectives on Approaches to Multi-agent Conversational AI
von: Clarke, Christopher, et al.
Veröffentlicht: (2024)