A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Fuente:
arXiv
Salvato in:
| Autori principali: | Ding, Peng, Kuang, Jun, Ma, Dan, Cao, Xuezhi, Xian, Yunsen, Chen, Jiajun, Huang, Shujian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
di: Ding, Peng, et al.
Pubblicazione: (2025)
di: Ding, Peng, et al.
Pubblicazione: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
di: Ding, Peng, et al.
Pubblicazione: (2024)
di: Ding, Peng, et al.
Pubblicazione: (2024)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
di: Ding, Peng, et al.
Pubblicazione: (2025)
di: Ding, Peng, et al.
Pubblicazione: (2025)
Electronic Resources: A Wolf in Sheep's Clothing?
di: Schaffner, Bradley L.
Pubblicazione: (2001)
di: Schaffner, Bradley L.
Pubblicazione: (2001)
The Wolf in Sheep’s Clothing: The Matthew Effect in Online Education
di: Amany Saleh
Pubblicazione: (2014)
di: Amany Saleh
Pubblicazione: (2014)
Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization
di: Cooper, Portia, et al.
Pubblicazione: (2024)
di: Cooper, Portia, et al.
Pubblicazione: (2024)
Lean Metabolic Dysfunction‐Associated Steatotic Liver Disease: A Wolf in Sheep's Clothing
di: Xixi Fang, et al.
Pubblicazione: (2025)
di: Xixi Fang, et al.
Pubblicazione: (2025)
A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG
di: Mu, Junjie, et al.
Pubblicazione: (2026)
di: Mu, Junjie, et al.
Pubblicazione: (2026)
Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
Acanthamoeba Encephalitis Presenting as Rapidly Developing Parkinsonism—A Wolf in Sheep's Clothing
di: Jacky Ganguly, et al.
Pubblicazione: (2025)
di: Jacky Ganguly, et al.
Pubblicazione: (2025)
A Wolf in Sheep’s Clothing: Extensive Musculoskeletal and Cutaneous TB Masquerading as Primary Erythema Nodosum
di: Tanner Shull, et al.
Pubblicazione: (2025)
di: Tanner Shull, et al.
Pubblicazione: (2025)
Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge
di: Li, Jiahuan, et al.
Pubblicazione: (2024)
di: Li, Jiahuan, et al.
Pubblicazione: (2024)
Large Language Models are Limited in Out-of-Context Knowledge Reasoning
di: Hu, Peng, et al.
Pubblicazione: (2024)
di: Hu, Peng, et al.
Pubblicazione: (2024)
SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
di: Ding, Peng, et al.
Pubblicazione: (2025)
di: Ding, Peng, et al.
Pubblicazione: (2025)
A Wolf in Sheep's Clothing: Practical Black-box Adversarial Attacks for Evading Learning-based Windows Malware Detection in the Wild
di: Ling, Xiang, et al.
Pubblicazione: (2024)
di: Ling, Xiang, et al.
Pubblicazione: (2024)
GPT in Sheep's Clothing: The Risk of Customized GPTs
di: Antebi, Sagiv, et al.
Pubblicazione: (2024)
di: Antebi, Sagiv, et al.
Pubblicazione: (2024)
Measuring Meaning Composition in the Human Brain with Composition Scores from Large Language Models
di: Gao, Changjiang, et al.
Pubblicazione: (2024)
di: Gao, Changjiang, et al.
Pubblicazione: (2024)
MT-PATCHER: Selective and Extendable Knowledge Distillation from Large Language Models for Machine Translation
di: Li, Jiahuan, et al.
Pubblicazione: (2024)
di: Li, Jiahuan, et al.
Pubblicazione: (2024)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
di: Yao, Yang, et al.
Pubblicazione: (2025)
di: Yao, Yang, et al.
Pubblicazione: (2025)
Exploring the Factual Consistency in Dialogue Comprehension of Large Language Models
di: She, Shuaijie, et al.
Pubblicazione: (2023)
di: She, Shuaijie, et al.
Pubblicazione: (2023)
Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation
di: Huang, Xu, et al.
Pubblicazione: (2024)
di: Huang, Xu, et al.
Pubblicazione: (2024)
Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions
di: Li, Jiahuan, et al.
Pubblicazione: (2023)
di: Li, Jiahuan, et al.
Pubblicazione: (2023)
Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
di: Hao, Run, et al.
Pubblicazione: (2025)
di: Hao, Run, et al.
Pubblicazione: (2025)
EDT: Improving Large Language Models' Generation by Entropy-based Dynamic Temperature Sampling
di: Zhang, Shimao, et al.
Pubblicazione: (2024)
di: Zhang, Shimao, et al.
Pubblicazione: (2024)
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
di: Jeong, Joonhyun, et al.
Pubblicazione: (2025)
di: Jeong, Joonhyun, et al.
Pubblicazione: (2025)
The Tower of Babel Revisited: Multilingual Jailbreak Prompts on Closed-Source Large Language Models
di: Huang, Linghan, et al.
Pubblicazione: (2025)
di: Huang, Linghan, et al.
Pubblicazione: (2025)
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
di: Yu, Jiahao, et al.
Pubblicazione: (2023)
di: Yu, Jiahao, et al.
Pubblicazione: (2023)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
di: Liu, Xiaogeng, et al.
Pubblicazione: (2023)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2023)
EvoJail: Evolutionary Diverse Jailbreak Prompt Generation for Large Language Models
di: Tang, Rui, et al.
Pubblicazione: (2026)
di: Tang, Rui, et al.
Pubblicazione: (2026)
On the Many Faces of Easily Covered Polytopes
di: Florentin, Dan I., et al.
Pubblicazione: (2024)
di: Florentin, Dan I., et al.
Pubblicazione: (2024)
Generalizing Dynamics Modeling More Easily from Representation Perspective
di: Wang, Yiming, et al.
Pubblicazione: (2026)
di: Wang, Yiming, et al.
Pubblicazione: (2026)
Extend Adversarial Policy Against Neural Machine Translation via Unknown Token
di: Zou, Wei, et al.
Pubblicazione: (2025)
di: Zou, Wei, et al.
Pubblicazione: (2025)
Beyond the Sequence: Statistics-Driven Pre-training for Stabilizing Sequential Recommendation Model
di: Wang, Sirui, et al.
Pubblicazione: (2024)
di: Wang, Sirui, et al.
Pubblicazione: (2024)
Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models
di: Ke, Shih-Wen, et al.
Pubblicazione: (2025)
di: Ke, Shih-Wen, et al.
Pubblicazione: (2025)
Trans-Zero: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
di: Zou, Wei, et al.
Pubblicazione: (2025)
di: Zou, Wei, et al.
Pubblicazione: (2025)
Exploiting Duality in Open Information Extraction with Predicate Prompt
di: Chen, Zhen, et al.
Pubblicazione: (2024)
di: Chen, Zhen, et al.
Pubblicazione: (2024)
StyleFool: Fooling Video Classification Systems via Style Transfer
di: Cao, Yuxin, et al.
Pubblicazione: (2022)
di: Cao, Yuxin, et al.
Pubblicazione: (2022)
Large Language Model Compression via the Nested Activation-Aware Decomposition
di: Lu, Jun, et al.
Pubblicazione: (2025)
di: Lu, Jun, et al.
Pubblicazione: (2025)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
di: Hu, Xiaomeng, et al.
Pubblicazione: (2024)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
di: Ding, Peng, et al.
Pubblicazione: (2025) -
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
di: Ding, Peng, et al.
Pubblicazione: (2024) -
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
di: Ding, Peng, et al.
Pubblicazione: (2025) -
Electronic Resources: A Wolf in Sheep's Clothing?
di: Schaffner, Bradley L.
Pubblicazione: (2001) -
The Wolf in Sheep’s Clothing: The Matthew Effect in Online Education
di: Amany Saleh
Pubblicazione: (2014)