Say It Differently: Linguistic Styles as Jailbreak Vectors
Fuente:
arXiv
Saved in:
| Main Authors: | Panda, Srikant, Rai, Avinash |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AccessEval: Benchmarking Disability Bias in Large Language Models
by: Panda, Srikant, et al.
Published: (2025)
by: Panda, Srikant, et al.
Published: (2025)
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
Jailbreaking to Jailbreak
by: Kritz, Jeremy, et al.
Published: (2025)
by: Kritz, Jeremy, et al.
Published: (2025)
A Multi-Task Role-Playing Agent Capable of Imitating Character Linguistic Styles
by: Chen, Siyuan, et al.
Published: (2024)
by: Chen, Siyuan, et al.
Published: (2024)
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
by: Panda, Srikant, et al.
Published: (2025)
by: Panda, Srikant, et al.
Published: (2025)
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
Support-Contra Asymmetry in LLM Explanations
by: Patil, Avinash
Published: (2025)
by: Patil, Avinash
Published: (2025)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
by: Zhou, Weikang, et al.
Published: (2024)
by: Zhou, Weikang, et al.
Published: (2024)
Many-Turn Jailbreaking
by: Yang, Xianjun, et al.
Published: (2025)
by: Yang, Xianjun, et al.
Published: (2025)
Embedding Style Beyond Topics: Analyzing Dispersion Effects Across Different Language Models
by: Icard, Benjamin, et al.
Published: (2025)
by: Icard, Benjamin, et al.
Published: (2025)
When Informal Text Breaks NLI: Tokenization Failure, Distribution Shift, and Targeted Mitigations
by: Aluguvelly, Avinash Goutham
Published: (2026)
by: Aluguvelly, Avinash Goutham
Published: (2026)
The Effectiveness of Style Vectors for Steering Large Language Models: A Human Evaluation
by: Diallo, Diaoulé, et al.
Published: (2026)
by: Diallo, Diaoulé, et al.
Published: (2026)
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
by: Gonçalves, Isabel, et al.
Published: (2025)
by: Gonçalves, Isabel, et al.
Published: (2025)
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
by: Ayoobi, Navid, et al.
Published: (2026)
by: Ayoobi, Navid, et al.
Published: (2026)
Chip-Tuning: Classify Before Language Models Say
by: Zhu, Fangwei, et al.
Published: (2024)
by: Zhu, Fangwei, et al.
Published: (2024)
FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
Advancing Reasoning in Large Language Models: Promising Methods and Approaches
by: Patil, Avinash, et al.
Published: (2025)
by: Patil, Avinash, et al.
Published: (2025)
Sarang at DEFACTIFY 4.0: Detecting AI-Generated Text Using Noised Data and an Ensemble of DeBERTa Models
by: Trivedi, Avinash, et al.
Published: (2025)
by: Trivedi, Avinash, et al.
Published: (2025)
Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
by: Camassa, Carolina, et al.
Published: (2026)
by: Camassa, Carolina, et al.
Published: (2026)
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
Playing Language Game with LLMs Leads to Jailbreaking
by: Peng, Yu, et al.
Published: (2024)
by: Peng, Yu, et al.
Published: (2024)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
by: Kabir, Md Rysul, et al.
Published: (2026)
by: Kabir, Md Rysul, et al.
Published: (2026)
Cognitive-Mental-LLM: Evaluating Reasoning in Large Language Models for Mental Health Prediction via Online Text
by: Patil, Avinash, et al.
Published: (2025)
by: Patil, Avinash, et al.
Published: (2025)
Saying the Unsaid: Revealing the Hidden Language of Multimodal Systems Through Telephone Games
by: Zhao, Juntu, et al.
Published: (2025)
by: Zhao, Juntu, et al.
Published: (2025)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
by: Puerto, Haritz, et al.
Published: (2026)
by: Puerto, Haritz, et al.
Published: (2026)
Exploring Chinese Humor Generation: A Study on Two-Part Allegorical Sayings
by: Xu, Rongwu
Published: (2024)
by: Xu, Rongwu
Published: (2024)
A Blueprint for AI-Driven Software Quality: Integrating LLMs with Established Standards
by: Patil, Avinash
Published: (2025)
by: Patil, Avinash
Published: (2025)
LLMs on a Budget? Say HOLA
by: Siddiqui, Zohaib Hasan, et al.
Published: (2025)
by: Siddiqui, Zohaib Hasan, et al.
Published: (2025)
Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
by: Weng, Zixuan, et al.
Published: (2025)
by: Weng, Zixuan, et al.
Published: (2025)
Poisoned LangChain: Jailbreak LLMs by LangChain
by: Wang, Ziqiu, et al.
Published: (2024)
by: Wang, Ziqiu, et al.
Published: (2024)
Cross-Lingual Jailbreak Detection via Semantic Codebooks
by: Alanova, Shirin, et al.
Published: (2026)
by: Alanova, Shirin, et al.
Published: (2026)
Merging Improves Self-Critique Against Jailbreak Attacks
by: Gallego, Victor
Published: (2024)
by: Gallego, Victor
Published: (2024)
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
Defending LLMs against Jailbreaking Attacks via Backtranslation
by: Wang, Yihan, et al.
Published: (2024)
by: Wang, Yihan, et al.
Published: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
The Art of Saying No: Contextual Noncompliance in Language Models
by: Brahman, Faeze, et al.
Published: (2024)
by: Brahman, Faeze, et al.
Published: (2024)
The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
by: Li, Yubo, et al.
Published: (2026)
by: Li, Yubo, et al.
Published: (2026)
The Cost of Thinking: Increased Jailbreak Risk in Large Language Models
by: Yang, Fan
Published: (2025)
by: Yang, Fan
Published: (2025)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
by: Giarrusso, Francesco, et al.
Published: (2025)
by: Giarrusso, Francesco, et al.
Published: (2025)
Similar Items
-
AccessEval: Benchmarking Disability Bias in Large Language Models
by: Panda, Srikant, et al.
Published: (2025) -
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
by: Joo, Seongho, et al.
Published: (2025) -
Jailbreaking to Jailbreak
by: Kritz, Jeremy, et al.
Published: (2025) -
A Multi-Task Role-Playing Agent Capable of Imitating Character Linguistic Styles
by: Chen, Siyuan, et al.
Published: (2024) -
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
by: Panda, Srikant, et al.
Published: (2025)