The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xing, Wang, Guanghui, Cui, Yanwei, Qiu, Wei, Li, Ziyuan, Zhu, Bing, He, Peiyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CloChat: Understanding How People Customize, Interact, and Experience Personas in Large Language Models
by: Ha, Juhye, et al.
Published: (2024)
by: Ha, Juhye, et al.
Published: (2024)
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
by: Fröhling, Leon, et al.
Published: (2024)
by: Fröhling, Leon, et al.
Published: (2024)
Direct Advantage Regression: Aligning LLMs with Online AI Reward
by: He, Li, et al.
Published: (2025)
by: He, Li, et al.
Published: (2025)
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties
by: Tang, Zixin, et al.
Published: (2025)
by: Tang, Zixin, et al.
Published: (2025)
Prompting ChatGPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts
by: He, Sui
Published: (2024)
by: He, Sui
Published: (2024)
Aligning LLMs with Individual Preferences via Interaction
by: Wu, Shujin, et al.
Published: (2024)
by: Wu, Shujin, et al.
Published: (2024)
Guardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
Dialogue Language Model with Large-Scale Persona Data Engineering
by: Hong, Mengze, et al.
Published: (2024)
by: Hong, Mengze, et al.
Published: (2024)
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
by: Liu, Naiming, et al.
Published: (2025)
by: Liu, Naiming, et al.
Published: (2025)
Aligning Language Models with Demonstrated Feedback
by: Shaikh, Omar, et al.
Published: (2024)
by: Shaikh, Omar, et al.
Published: (2024)
Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
by: Amin, Danial, et al.
Published: (2025)
by: Amin, Danial, et al.
Published: (2025)
CALYPSO: LLMs as Dungeon Masters' Assistants
by: Zhu, Andrew, et al.
Published: (2023)
by: Zhu, Andrew, et al.
Published: (2023)
LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs
by: Wu, Tongshuang, et al.
Published: (2023)
by: Wu, Tongshuang, et al.
Published: (2023)
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
by: Afzoon, Saleh, et al.
Published: (2026)
by: Afzoon, Saleh, et al.
Published: (2026)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
by: Afzoon, Saleh, et al.
Published: (2026)
by: Afzoon, Saleh, et al.
Published: (2026)
ReSpark: Leveraging Previous Data Reports as References to Generate New Reports with LLMs
by: Tian, Yuan, et al.
Published: (2025)
by: Tian, Yuan, et al.
Published: (2025)
Generating Educational Materials with Different Levels of Readability using LLMs
by: Huang, Chieh-Yang, et al.
Published: (2024)
by: Huang, Chieh-Yang, et al.
Published: (2024)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
by: Yu, Yeyong, et al.
Published: (2024)
by: Yu, Yeyong, et al.
Published: (2024)
SensorPersona: An LLM-Empowered System for Continual Persona Extraction from Longitudinal Mobile Sensor Streams
by: Yang, Bufang, et al.
Published: (2026)
by: Yang, Bufang, et al.
Published: (2026)
Customizing ChatGPT for Second Language Speaking Practice: Genuine Support or Just a Marketing Gimmick?
by: Meng, Fanfei
Published: (2026)
by: Meng, Fanfei
Published: (2026)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Aligning Stuttered-Speech Research with End-User Needs: Scoping Review, Survey, and Guidelines
by: Toyin, Hawau Olamide, et al.
Published: (2026)
by: Toyin, Hawau Olamide, et al.
Published: (2026)
Aligning Human-AI-Interaction Trust for Mental Health Support: Survey and Position for Multi-Stakeholders
by: Sun, Xin, et al.
Published: (2026)
by: Sun, Xin, et al.
Published: (2026)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
LLMs' ways of seeing User Personas
by: Panda, Swaroop
Published: (2024)
by: Panda, Swaroop
Published: (2024)
Comparing How a Chatbot References User Utterances from Previous Chatting Sessions: An Investigation of Users' Privacy Concerns and Perceptions
by: Cox, Samuel Rhys, et al.
Published: (2023)
by: Cox, Samuel Rhys, et al.
Published: (2023)
Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment
by: Kim, Yoonsu, et al.
Published: (2024)
by: Kim, Yoonsu, et al.
Published: (2024)
Towards Stable and Personalised Profiles for Lexical Alignment in Spoken Human-Agent Dialogue
by: Schaaij, Keara, et al.
Published: (2025)
by: Schaaij, Keara, et al.
Published: (2025)
Completing A Systematic Review in Hours instead of Months with Interactive AI Agents
by: Qiu, Rui, et al.
Published: (2025)
by: Qiu, Rui, et al.
Published: (2025)
After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions
by: Xuan, Ziyi, et al.
Published: (2026)
by: Xuan, Ziyi, et al.
Published: (2026)
Robots in the Middle: Evaluating LLMs in Dispute Resolution
by: Tan, Jinzhe, et al.
Published: (2024)
by: Tan, Jinzhe, et al.
Published: (2024)
Similar Items
-
CloChat: Understanding How People Customize, Interact, and Experience Personas in Large Language Models
by: Ha, Juhye, et al.
Published: (2024) -
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
by: Fröhling, Leon, et al.
Published: (2024) -
Direct Advantage Regression: Aligning LLMs with Online AI Reward
by: He, Li, et al.
Published: (2025) -
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
by: Wang, Ziyi, et al.
Published: (2025) -
Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties
by: Tang, Zixin, et al.
Published: (2025)