Persona Jailbreaking in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Sandhan, Jivnesh, Cheng, Fei, Sandhan, Tushar, Murawaki, Yugo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAPE: Context-Aware Personality Evaluation Framework for Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2025)
by: Sandhan, Jivnesh, et al.
Published: (2025)
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages
by: Ray, Pretam, et al.
Published: (2024)
by: Ray, Pretam, et al.
Published: (2024)
Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
by: Yan, Ruiyi, et al.
Published: (2025)
by: Yan, Ruiyi, et al.
Published: (2025)
Still Not There: Can LLMs Outperform Smaller Task-Specific Seq2Seq Models on the Poetry-to-Prose Conversion Task?
by: Das, Kunal Kingkar, et al.
Published: (2025)
by: Das, Kunal Kingkar, et al.
Published: (2025)
Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
by: Sarkar, Sujoy, et al.
Published: (2025)
by: Sarkar, Sujoy, et al.
Published: (2025)
Principal Component Analysis as a Sanity Check for Bayesian Phylolinguistic Reconstruction
by: Murawaki, Yugo
Published: (2024)
by: Murawaki, Yugo
Published: (2024)
Vision-Aided Online A* Path Planning for Efficient and Safe Navigation of Service Robots
by: Kumar, Praveen, et al.
Published: (2025)
by: Kumar, Praveen, et al.
Published: (2025)
Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
by: Zhong, Chengzhi, et al.
Published: (2025)
by: Zhong, Chengzhi, et al.
Published: (2025)
Efficient Provably Secure Linguistic Steganography via Range Coding
by: Yan, Ruiyi, et al.
Published: (2026)
by: Yan, Ruiyi, et al.
Published: (2026)
How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations
by: Takenami, Yoshiki, et al.
Published: (2025)
by: Takenami, Yoshiki, et al.
Published: (2025)
Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography
by: Yan, Ruiyi, et al.
Published: (2026)
by: Yan, Ruiyi, et al.
Published: (2026)
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
by: Matta, Shiho, et al.
Published: (2024)
by: Matta, Shiho, et al.
Published: (2024)
Beyond English-Centric LLMs: What Language Do Multilingual Language Models Think in?
by: Zhong, Chengzhi, et al.
Published: (2024)
by: Zhong, Chengzhi, et al.
Published: (2024)
Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking
by: Xu, Nan, et al.
Published: (2023)
by: Xu, Nan, et al.
Published: (2023)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
by: Zhang, Zhexin, et al.
Published: (2023)
by: Zhang, Zhexin, et al.
Published: (2023)
Jailbreaking Large Language Models with Morality Attacks
by: Su, Ying, et al.
Published: (2026)
by: Su, Ying, et al.
Published: (2026)
Weak-to-Strong Jailbreaking on Large Language Models
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
Diversity Helps Jailbreak Large Language Models
by: Zhao, Weiliang, et al.
Published: (2024)
by: Zhao, Weiliang, et al.
Published: (2024)
Multilingual Jailbreak Challenges in Large Language Models
by: Deng, Yue, et al.
Published: (2023)
by: Deng, Yue, et al.
Published: (2023)
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
by: Mei, Lingrui, et al.
Published: (2024)
by: Mei, Lingrui, et al.
Published: (2024)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
by: Zhou, Weikang, et al.
Published: (2024)
by: Zhou, Weikang, et al.
Published: (2024)
AJF: Adaptive Jailbreak Framework Based on the Comprehension Ability of Black-Box Large Language Models
by: Yu, Mingyu, et al.
Published: (2025)
by: Yu, Mingyu, et al.
Published: (2025)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
Structured Semantic Cloaking for Jailbreak Attacks on Large Language Models
by: Sun, Xiaobing, et al.
Published: (2026)
by: Sun, Xiaobing, et al.
Published: (2026)
Evaluating Large Language Model Biases in Persona-Steered Generation
by: Liu, Andy, et al.
Published: (2024)
by: Liu, Andy, et al.
Published: (2024)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Can Large Language Models Automatically Jailbreak GPT-4V?
by: Wu, Yuanwei, et al.
Published: (2024)
by: Wu, Yuanwei, et al.
Published: (2024)
Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models
by: Li, Jiahui, et al.
Published: (2024)
by: Li, Jiahui, et al.
Published: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
by: Lee, Isack, et al.
Published: (2024)
by: Lee, Isack, et al.
Published: (2024)
JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
by: Jin, Haibo, et al.
Published: (2024)
by: Jin, Haibo, et al.
Published: (2024)
Imperceptible Jailbreaking against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2025)
by: Gao, Kuofeng, et al.
Published: (2025)
Dialogue Language Model with Large-Scale Persona Data Engineering
by: Hong, Mengze, et al.
Published: (2024)
by: Hong, Mengze, et al.
Published: (2024)
The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models
by: Xiao, Yunze, et al.
Published: (2026)
by: Xiao, Yunze, et al.
Published: (2026)
Beyond Static Personas: Situational Personality Steering for Large Language Models
by: Wei, Zesheng, et al.
Published: (2026)
by: Wei, Zesheng, et al.
Published: (2026)
LASH: Adaptive Semantic Hybridization for Black-Box Jailbreaking of Large Language Models
by: Nafi, Abdullah Al Nomaan, et al.
Published: (2026)
by: Nafi, Abdullah Al Nomaan, et al.
Published: (2026)
Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge
by: Lu, Weikai, et al.
Published: (2024)
by: Lu, Weikai, et al.
Published: (2024)
Similar Items
-
CAPE: Context-Aware Personality Evaluation Framework for Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2025) -
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026) -
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages
by: Ray, Pretam, et al.
Published: (2024) -
Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
by: Yan, Ruiyi, et al.
Published: (2025) -
Still Not There: Can LLMs Outperform Smaller Task-Specific Seq2Seq Models on the Poetry-to-Prose Conversion Task?
by: Das, Kunal Kingkar, et al.
Published: (2025)