Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | de Araujo, Pedro Henrique Luz, Hedderich, Michael A., Modarressi, Ali, Schuetze, Hinrich, Roth, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
by: Fayyaz, Mohsen, et al.
Published: (2025)
by: Fayyaz, Mohsen, et al.
Published: (2025)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
Consistent Document-Level Relation Extraction via Counterfactuals
by: Modarressi, Ali, et al.
Published: (2024)
by: Modarressi, Ali, et al.
Published: (2024)
Helpful assistant or fruitful facilitator? Investigating how personas affect language model behavior
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2024)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2024)
Functionality learning through specification instructions
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2023)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2023)
RET-LLM: Towards a General Read-Write Memory for Large Language Models
by: Modarressi, Ali, et al.
Published: (2023)
by: Modarressi, Ali, et al.
Published: (2023)
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
by: Hakimi, Ahmad Dawar, et al.
Published: (2025)
by: Hakimi, Ahmad Dawar, et al.
Published: (2025)
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
by: Modarressi, Ali, et al.
Published: (2024)
by: Modarressi, Ali, et al.
Published: (2024)
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge
by: Taghavi, Zeinab Sadat, et al.
Published: (2025)
by: Taghavi, Zeinab Sadat, et al.
Published: (2025)
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles
by: Xia, Yuxi, et al.
Published: (2025)
by: Xia, Yuxi, et al.
Published: (2025)
RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-Following
by: Lu, Junru, et al.
Published: (2025)
by: Lu, Junru, et al.
Published: (2025)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
by: Zhao, Raoyuan, et al.
Published: (2026)
by: Zhao, Raoyuan, et al.
Published: (2026)
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
by: Liu, Yuxin, et al.
Published: (2026)
by: Liu, Yuxin, et al.
Published: (2026)
Enhancing Persona Consistency for LLMs' Role-Playing using Persona-Aware Contrastive Learning
by: Ji, Ke, et al.
Published: (2025)
by: Ji, Ke, et al.
Published: (2025)
NoLiMa: Long-Context Evaluation Beyond Literal Matching
by: Modarressi, Ali, et al.
Published: (2025)
by: Modarressi, Ali, et al.
Published: (2025)
MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
by: Tseng, Yu-Min, et al.
Published: (2024)
by: Tseng, Yu-Min, et al.
Published: (2024)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
by: Zhou, Lingfeng, et al.
Published: (2025)
by: Zhou, Lingfeng, et al.
Published: (2025)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Exploring prompts to elicit memorization in masked language model-based named entity recognition
by: Xia, Yuxi, et al.
Published: (2024)
by: Xia, Yuxi, et al.
Published: (2024)
CharacterGPT: A Persona Reconstruction Framework for Role-Playing Agents
by: Park, Jeiyoon, et al.
Published: (2024)
by: Park, Jeiyoon, et al.
Published: (2024)
From Persona to Personalization: A Survey on Role-Playing Language Agents
by: Chen, Jiangjie, et al.
Published: (2024)
by: Chen, Jiangjie, et al.
Published: (2024)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
by: Wang, Xiaoyang, et al.
Published: (2025)
by: Wang, Xiaoyang, et al.
Published: (2025)
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework
by: Yang, Bohao, et al.
Published: (2024)
by: Yang, Bohao, et al.
Published: (2024)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
by: Costa, Davi Bastos, et al.
Published: (2025)
by: Costa, Davi Bastos, et al.
Published: (2025)
Steering MoE LLMs via Expert (De)Activation
by: Fayyaz, Mohsen, et al.
Published: (2025)
by: Fayyaz, Mohsen, et al.
Published: (2025)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs
by: Tang, Wenqiu, et al.
Published: (2026)
by: Tang, Wenqiu, et al.
Published: (2026)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
by: Zhang, Jiaqiao, et al.
Published: (2026)
by: Zhang, Jiaqiao, et al.
Published: (2026)
From Role-Play to Drama-Interaction: An LLM Solution
by: Wu, Weiqi, et al.
Published: (2024)
by: Wu, Weiqi, et al.
Published: (2024)
When Contextual Inference Fails: Cancelability in Interactive Instruction Following
by: Bila, Natalia, et al.
Published: (2026)
by: Bila, Natalia, et al.
Published: (2026)
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
by: Bianchi, Federico, et al.
Published: (2023)
by: Bianchi, Federico, et al.
Published: (2023)
Safety Training Persists Through Helpfulness Optimization in LLM Agents
by: Plaut, Benjamin
Published: (2026)
by: Plaut, Benjamin
Published: (2026)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Can LLMs Simulate Personas with Reversed Performance? A Systematic Investigation for Counterfactual Instruction Following in Math Reasoning Context
by: Kumar, Sai Adith Senthil, et al.
Published: (2025)
by: Kumar, Sai Adith Senthil, et al.
Published: (2025)
Similar Items
-
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
by: Fayyaz, Mohsen, et al.
Published: (2025) -
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
by: Zhao, Raoyuan, et al.
Published: (2025) -
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025) -
Consistent Document-Level Relation Extraction via Counterfactuals
by: Modarressi, Ali, et al.
Published: (2024) -
Helpful assistant or fruitful facilitator? Investigating how personas affect language model behavior
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2024)