A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ding, Peng, Kuang, Jun, Ma, Dan, Cao, Xuezhi, Xian, Yunsen, Chen, Jiajun, Huang, Shujian |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
par: Ding, Peng, et autres
Publié: (2025)
par: Ding, Peng, et autres
Publié: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
par: Ding, Peng, et autres
Publié: (2024)
par: Ding, Peng, et autres
Publié: (2024)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
par: Ding, Peng, et autres
Publié: (2025)
par: Ding, Peng, et autres
Publié: (2025)
Electronic Resources: A Wolf in Sheep's Clothing?
par: Schaffner, Bradley L.
Publié: (2001)
par: Schaffner, Bradley L.
Publié: (2001)
The Wolf in Sheep’s Clothing: The Matthew Effect in Online Education
par: Amany Saleh
Publié: (2014)
par: Amany Saleh
Publié: (2014)
Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization
par: Cooper, Portia, et autres
Publié: (2024)
par: Cooper, Portia, et autres
Publié: (2024)
Lean Metabolic Dysfunction‐Associated Steatotic Liver Disease: A Wolf in Sheep's Clothing
par: Xixi Fang, et autres
Publié: (2025)
par: Xixi Fang, et autres
Publié: (2025)
A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG
par: Mu, Junjie, et autres
Publié: (2026)
par: Mu, Junjie, et autres
Publié: (2026)
Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
par: Guo, Xuyang, et autres
Publié: (2025)
par: Guo, Xuyang, et autres
Publié: (2025)
Acanthamoeba Encephalitis Presenting as Rapidly Developing Parkinsonism—A Wolf in Sheep's Clothing
par: Jacky Ganguly, et autres
Publié: (2025)
par: Jacky Ganguly, et autres
Publié: (2025)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
par: Saiem, Bijoy Ahmed, et autres
Publié: (2024)
par: Saiem, Bijoy Ahmed, et autres
Publié: (2024)
A Wolf in Sheep’s Clothing: Extensive Musculoskeletal and Cutaneous TB Masquerading as Primary Erythema Nodosum
par: Tanner Shull, et autres
Publié: (2025)
par: Tanner Shull, et autres
Publié: (2025)
Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge
par: Li, Jiahuan, et autres
Publié: (2024)
par: Li, Jiahuan, et autres
Publié: (2024)
Large Language Models are Limited in Out-of-Context Knowledge Reasoning
par: Hu, Peng, et autres
Publié: (2024)
par: Hu, Peng, et autres
Publié: (2024)
SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
par: Ding, Peng, et autres
Publié: (2025)
par: Ding, Peng, et autres
Publié: (2025)
A Wolf in Sheep's Clothing: Practical Black-box Adversarial Attacks for Evading Learning-based Windows Malware Detection in the Wild
par: Ling, Xiang, et autres
Publié: (2024)
par: Ling, Xiang, et autres
Publié: (2024)
GPT in Sheep's Clothing: The Risk of Customized GPTs
par: Antebi, Sagiv, et autres
Publié: (2024)
par: Antebi, Sagiv, et autres
Publié: (2024)
Measuring Meaning Composition in the Human Brain with Composition Scores from Large Language Models
par: Gao, Changjiang, et autres
Publié: (2024)
par: Gao, Changjiang, et autres
Publié: (2024)
MT-PATCHER: Selective and Extendable Knowledge Distillation from Large Language Models for Machine Translation
par: Li, Jiahuan, et autres
Publié: (2024)
par: Li, Jiahuan, et autres
Publié: (2024)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
par: Yao, Yang, et autres
Publié: (2025)
par: Yao, Yang, et autres
Publié: (2025)
Exploring the Factual Consistency in Dialogue Comprehension of Large Language Models
par: She, Shuaijie, et autres
Publié: (2023)
par: She, Shuaijie, et autres
Publié: (2023)
Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation
par: Huang, Xu, et autres
Publié: (2024)
par: Huang, Xu, et autres
Publié: (2024)
Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions
par: Li, Jiahuan, et autres
Publié: (2023)
par: Li, Jiahuan, et autres
Publié: (2023)
Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
par: Hao, Run, et autres
Publié: (2025)
par: Hao, Run, et autres
Publié: (2025)
EDT: Improving Large Language Models' Generation by Entropy-based Dynamic Temperature Sampling
par: Zhang, Shimao, et autres
Publié: (2024)
par: Zhang, Shimao, et autres
Publié: (2024)
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
par: Jeong, Joonhyun, et autres
Publié: (2025)
par: Jeong, Joonhyun, et autres
Publié: (2025)
The Tower of Babel Revisited: Multilingual Jailbreak Prompts on Closed-Source Large Language Models
par: Huang, Linghan, et autres
Publié: (2025)
par: Huang, Linghan, et autres
Publié: (2025)
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
par: Yu, Jiahao, et autres
Publié: (2023)
par: Yu, Jiahao, et autres
Publié: (2023)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
par: Liu, Xiaogeng, et autres
Publié: (2023)
par: Liu, Xiaogeng, et autres
Publié: (2023)
EvoJail: Evolutionary Diverse Jailbreak Prompt Generation for Large Language Models
par: Tang, Rui, et autres
Publié: (2026)
par: Tang, Rui, et autres
Publié: (2026)
On the Many Faces of Easily Covered Polytopes
par: Florentin, Dan I., et autres
Publié: (2024)
par: Florentin, Dan I., et autres
Publié: (2024)
Generalizing Dynamics Modeling More Easily from Representation Perspective
par: Wang, Yiming, et autres
Publié: (2026)
par: Wang, Yiming, et autres
Publié: (2026)
Extend Adversarial Policy Against Neural Machine Translation via Unknown Token
par: Zou, Wei, et autres
Publié: (2025)
par: Zou, Wei, et autres
Publié: (2025)
Beyond the Sequence: Statistics-Driven Pre-training for Stabilizing Sequential Recommendation Model
par: Wang, Sirui, et autres
Publié: (2024)
par: Wang, Sirui, et autres
Publié: (2024)
Trans-Zero: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
par: Zou, Wei, et autres
Publié: (2025)
par: Zou, Wei, et autres
Publié: (2025)
Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models
par: Ke, Shih-Wen, et autres
Publié: (2025)
par: Ke, Shih-Wen, et autres
Publié: (2025)
Exploiting Duality in Open Information Extraction with Predicate Prompt
par: Chen, Zhen, et autres
Publié: (2024)
par: Chen, Zhen, et autres
Publié: (2024)
StyleFool: Fooling Video Classification Systems via Style Transfer
par: Cao, Yuxin, et autres
Publié: (2022)
par: Cao, Yuxin, et autres
Publié: (2022)
Large Language Model Compression via the Nested Activation-Aware Decomposition
par: Lu, Jun, et autres
Publié: (2025)
par: Lu, Jun, et autres
Publié: (2025)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
par: Hu, Xiaomeng, et autres
Publié: (2024)
par: Hu, Xiaomeng, et autres
Publié: (2024)
Documents similaires
-
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
par: Ding, Peng, et autres
Publié: (2025) -
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
par: Ding, Peng, et autres
Publié: (2024) -
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
par: Ding, Peng, et autres
Publié: (2025) -
Electronic Resources: A Wolf in Sheep's Clothing?
par: Schaffner, Bradley L.
Publié: (2001) -
The Wolf in Sheep’s Clothing: The Matthew Effect in Online Education
par: Amany Saleh
Publié: (2014)