Bullying the Machine: How Personas Increase LLM Vulnerability
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Ziwei, Sanghi, Udit, Kankanhalli, Mohan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hallucination is Inevitable: An Innate Limitation of Large Language Models
by: Xu, Ziwei, et al.
Published: (2024)
by: Xu, Ziwei, et al.
Published: (2024)
Reasoning LLMs are Wandering Solution Explorers
by: Lu, Jiahao, et al.
Published: (2025)
by: Lu, Jiahao, et al.
Published: (2025)
How to Understand Named Entities: Using Common Sense for News Captioning
by: Xu, Ning, et al.
Published: (2024)
by: Xu, Ning, et al.
Published: (2024)
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
by: Xu, Danni, et al.
Published: (2025)
by: Xu, Danni, et al.
Published: (2025)
Strong Preferences Affect the Robustness of Preference Models and Value Alignment
by: Xu, Ziwei, et al.
Published: (2024)
by: Xu, Ziwei, et al.
Published: (2024)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
by: Nawal, Aditya, et al.
Published: (2026)
by: Nawal, Aditya, et al.
Published: (2026)
Nine Ways to Break Copyright Law and Why Our LLM Won't: A Fair Use Aligned Generation Framework
by: Sharma, Aakash Sen, et al.
Published: (2025)
by: Sharma, Aakash Sen, et al.
Published: (2025)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
by: Kim, Jiseon, et al.
Published: (2025)
by: Kim, Jiseon, et al.
Published: (2025)
Simulating Misinformation Vulnerabilities With Agent Personas
by: Farr, David, et al.
Published: (2025)
by: Farr, David, et al.
Published: (2025)
Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
by: Fan, Hehe, et al.
Published: (2025)
by: Fan, Hehe, et al.
Published: (2025)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Tracing Persona Vectors Through LLM Pretraining
by: Moskvoretskii, Viktor, et al.
Published: (2026)
by: Moskvoretskii, Viktor, et al.
Published: (2026)
Where is the Mind? Persona Vectors and LLM Individuation
by: Beckmann, Pierre, et al.
Published: (2026)
by: Beckmann, Pierre, et al.
Published: (2026)
Polypersona: Persona-Grounded LLM for Synthetic Survey Responses
by: Dash, Tejaswani, et al.
Published: (2025)
by: Dash, Tejaswani, et al.
Published: (2025)
SensorPersona: An LLM-Empowered System for Continual Persona Extraction from Longitudinal Mobile Sensor Streams
by: Yang, Bufang, et al.
Published: (2026)
by: Yang, Bufang, et al.
Published: (2026)
SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection
by: Kazemi, Arefeh, et al.
Published: (2025)
by: Kazemi, Arefeh, et al.
Published: (2025)
A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities
by: Chen, Jiaqi, et al.
Published: (2026)
by: Chen, Jiaqi, et al.
Published: (2026)
Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication
by: Rizvi, Naba, et al.
Published: (2026)
by: Rizvi, Naba, et al.
Published: (2026)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
by: Singh, Himanshu, et al.
Published: (2026)
by: Singh, Himanshu, et al.
Published: (2026)
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
by: Choi, Junhyuk, et al.
Published: (2025)
by: Choi, Junhyuk, et al.
Published: (2025)
TalkDep: Clinically Grounded LLM Personas for Conversation-Centric Depression Screening
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
by: Rupprecht, Jens, et al.
Published: (2025)
by: Rupprecht, Jens, et al.
Published: (2025)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
Not All Personas Are Worth It: Culture-Reflective Persona Data Augmentation
by: Han, Ji-Eun, et al.
Published: (2025)
by: Han, Ji-Eun, et al.
Published: (2025)
Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding
by: Borchers, Conrad, et al.
Published: (2025)
by: Borchers, Conrad, et al.
Published: (2025)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
by: Wang, Xintao, et al.
Published: (2025)
by: Wang, Xintao, et al.
Published: (2025)
PersonaMatrix: A Recipe for Persona-Aware Evaluation of Legal Summarization
by: Pang, Tsz Fung, et al.
Published: (2025)
by: Pang, Tsz Fung, et al.
Published: (2025)
The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models
by: Abdullahi, Tassallah, et al.
Published: (2026)
by: Abdullahi, Tassallah, et al.
Published: (2026)
$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales
by: Cao, Jun, et al.
Published: (2025)
by: Cao, Jun, et al.
Published: (2025)
Increasing the Robustness of the Fine-tuned Multilingual Machine-Generated Text Detectors
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
Population-Aligned Persona Generation for LLM-based Social Simulation
by: Hu, Zhengyu, et al.
Published: (2025)
by: Hu, Zhengyu, et al.
Published: (2025)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
by: Das, Nilanjana, et al.
Published: (2024)
by: Das, Nilanjana, et al.
Published: (2024)
Localizing Persona Representations in LLMs
by: Cintas, Celia, et al.
Published: (2025)
by: Cintas, Celia, et al.
Published: (2025)
MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
by: Lee, Kyungro, et al.
Published: (2025)
by: Lee, Kyungro, et al.
Published: (2025)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
by: Mo, Lingbo, et al.
Published: (2023)
by: Mo, Lingbo, et al.
Published: (2023)
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
by: Zugecova, Aneta, et al.
Published: (2024)
by: Zugecova, Aneta, et al.
Published: (2024)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Similar Items
-
Hallucination is Inevitable: An Innate Limitation of Large Language Models
by: Xu, Ziwei, et al.
Published: (2024) -
Reasoning LLMs are Wandering Solution Explorers
by: Lu, Jiahao, et al.
Published: (2025) -
How to Understand Named Entities: Using Common Sense for News Captioning
by: Xu, Ning, et al.
Published: (2024) -
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
by: Xu, Danni, et al.
Published: (2025) -
Strong Preferences Affect the Robustness of Preference Models and Value Alignment
by: Xu, Ziwei, et al.
Published: (2024)