EnJa: Ensemble Jailbreak on Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Jiahao, Wang, Zilong, Wang, Ruofan, Ma, Xingjun, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
par: Yao, Yang, et autres
Publié: (2025)
par: Yao, Yang, et autres
Publié: (2025)
Imperceptible Jailbreaking against Large Language Models
par: Gao, Kuofeng, et autres
Publié: (2025)
par: Gao, Kuofeng, et autres
Publié: (2025)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
par: Hua, Peichun, et autres
Publié: (2025)
par: Hua, Peichun, et autres
Publié: (2025)
Jailbreaking Large Language Models with Symbolic Mathematics
par: Bethany, Emet, et autres
Publié: (2024)
par: Bethany, Emet, et autres
Publié: (2024)
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
par: Bai, Yang, et autres
Publié: (2024)
par: Bai, Yang, et autres
Publié: (2024)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
par: Lin, Shi, et autres
Publié: (2024)
par: Lin, Shi, et autres
Publié: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
par: Hu, Xiaomeng, et autres
Publié: (2024)
par: Hu, Xiaomeng, et autres
Publié: (2024)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
par: Wei, Zeming, et autres
Publié: (2023)
par: Wei, Zeming, et autres
Publié: (2023)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
par: Boreiko, Valentyn, et autres
Publié: (2024)
par: Boreiko, Valentyn, et autres
Publié: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
par: Yi, Sibo, et autres
Publié: (2024)
par: Yi, Sibo, et autres
Publié: (2024)
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
par: Goel, Aman, et autres
Publié: (2025)
par: Goel, Aman, et autres
Publié: (2025)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
par: Li, Xiao, et autres
Publié: (2024)
par: Li, Xiao, et autres
Publié: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
par: Reddy, Aashray, et autres
Publié: (2025)
par: Reddy, Aashray, et autres
Publié: (2025)
Rethinking How to Evaluate Language Model Jailbreak
par: Cai, Hongyu, et autres
Publié: (2024)
par: Cai, Hongyu, et autres
Publié: (2024)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
par: Zheng, Xiaosen, et autres
Publié: (2024)
par: Zheng, Xiaosen, et autres
Publié: (2024)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
par: Peng, Benji, et autres
Publié: (2024)
par: Peng, Benji, et autres
Publié: (2024)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
par: Zhao, Andrew, et autres
Publié: (2024)
par: Zhao, Andrew, et autres
Publié: (2024)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
par: Fang, Zheng, et autres
Publié: (2026)
par: Fang, Zheng, et autres
Publié: (2026)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
par: Saiem, Bijoy Ahmed, et autres
Publié: (2024)
par: Saiem, Bijoy Ahmed, et autres
Publié: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
par: Tu, Shangqing, et autres
Publié: (2024)
par: Tu, Shangqing, et autres
Publié: (2024)
Instructional Fingerprinting of Large Language Models
par: Xu, Jiashu, et autres
Publié: (2024)
par: Xu, Jiashu, et autres
Publié: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
par: Chu, Junjie, et autres
Publié: (2024)
par: Chu, Junjie, et autres
Publié: (2024)
Low-Resource Languages Jailbreak GPT-4
par: Yong, Zheng-Xin, et autres
Publié: (2023)
par: Yong, Zheng-Xin, et autres
Publié: (2023)
Jailbreaking with Universal Multi-Prompts
par: Hsu, Yu-Ling, et autres
Publié: (2025)
par: Hsu, Yu-Ling, et autres
Publié: (2025)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
par: Mo, Yichuan, et autres
Publié: (2024)
par: Mo, Yichuan, et autres
Publié: (2024)
Internal Safety Collapse in Frontier Large Language Models
par: Wu, Yutao, et autres
Publié: (2026)
par: Wu, Yutao, et autres
Publié: (2026)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
par: Wang, Zhepeng, et autres
Publié: (2024)
par: Wang, Zhepeng, et autres
Publié: (2024)
On the Role of Attention Heads in Large Language Model Safety
par: Zhou, Zhenhong, et autres
Publié: (2024)
par: Zhou, Zhenhong, et autres
Publié: (2024)
Detecting Training Data of Large Language Models via Expectation Maximization
par: Kim, Gyuwan, et autres
Publié: (2024)
par: Kim, Gyuwan, et autres
Publié: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
par: Xu, Jiashu, et autres
Publié: (2023)
par: Xu, Jiashu, et autres
Publié: (2023)
Jailbreaking in the Haystack
par: Shah, Rishi Rajesh, et autres
Publié: (2025)
par: Shah, Rishi Rajesh, et autres
Publié: (2025)
DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models
par: Sun, Ye, et autres
Publié: (2026)
par: Sun, Ye, et autres
Publié: (2026)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
par: Jiang, Weisen, et autres
Publié: (2025)
par: Jiang, Weisen, et autres
Publié: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
par: Chen, Taiye, et autres
Publié: (2025)
par: Chen, Taiye, et autres
Publié: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
par: Hasan, Adib, et autres
Publié: (2024)
par: Hasan, Adib, et autres
Publié: (2024)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
par: Gu, Tianle, et autres
Publié: (2025)
par: Gu, Tianle, et autres
Publié: (2025)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
par: Galinkin, Erick, et autres
Publié: (2024)
par: Galinkin, Erick, et autres
Publié: (2024)
Jailbreaking LLMs via Calibration
par: Lu, Yuxuan, et autres
Publié: (2026)
par: Lu, Yuxuan, et autres
Publié: (2026)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
par: Wang, Zilong, et autres
Publié: (2025)
par: Wang, Zilong, et autres
Publié: (2025)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
par: Li, Lijun, et autres
Publié: (2024)
par: Li, Lijun, et autres
Publié: (2024)
Documents similaires
-
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
par: Yao, Yang, et autres
Publié: (2025) -
Imperceptible Jailbreaking against Large Language Models
par: Gao, Kuofeng, et autres
Publié: (2025) -
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
par: Hua, Peichun, et autres
Publié: (2025) -
Jailbreaking Large Language Models with Symbolic Mathematics
par: Bethany, Emet, et autres
Publié: (2024) -
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
par: Bai, Yang, et autres
Publié: (2024)