Rethinking How to Evaluate Language Model Jailbreak
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Hongyu, Arunasalam, Arjun, Lin, Leo Y., Bianchi, Antonio, Celik, Z. Berkay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring and Developing a Pre-Model Safeguard with Draft Models
von: Cai, Hongyu, et al.
Veröffentlicht: (2026)
von: Cai, Hongyu, et al.
Veröffentlicht: (2026)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
von: Hua, Peichun, et al.
Veröffentlicht: (2025)
von: Hua, Peichun, et al.
Veröffentlicht: (2025)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
von: Lin, Shi, et al.
Veröffentlicht: (2024)
von: Lin, Shi, et al.
Veröffentlicht: (2024)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
EnJa: Ensemble Jailbreak on Large Language Models
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
von: Wei, Zeming, et al.
Veröffentlicht: (2023)
von: Wei, Zeming, et al.
Veröffentlicht: (2023)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Low-Resource Languages Jailbreak GPT-4
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2023)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2023)
International Students and Scams: At Risk Abroad
von: Zhang, Katherine, et al.
Veröffentlicht: (2025)
von: Zhang, Katherine, et al.
Veröffentlicht: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
Jailbreaking in the Haystack
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
von: Li, Xiao, et al.
Veröffentlicht: (2024)
von: Li, Xiao, et al.
Veröffentlicht: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
von: Goel, Aman, et al.
Veröffentlicht: (2025)
von: Goel, Aman, et al.
Veröffentlicht: (2025)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
Jailbreaking with Universal Multi-Prompts
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
von: Hsu, Yu-Ling, et al.
Veröffentlicht: (2025)
Jailbreaking LLMs via Calibration
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
von: Yao, Yang, et al.
Veröffentlicht: (2025)
von: Yao, Yang, et al.
Veröffentlicht: (2025)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
von: Peng, Benji, et al.
Veröffentlicht: (2024)
von: Peng, Benji, et al.
Veröffentlicht: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
Universal Jailbreak Backdoors from Poisoned Human Feedback
von: Rando, Javier, et al.
Veröffentlicht: (2023)
von: Rando, Javier, et al.
Veröffentlicht: (2023)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Fang, Zhicheng, et al.
Veröffentlicht: (2026)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
Enhancing LLM-based Autonomous Driving Agents to Mitigate Perception Attacks
von: Song, Ruoyu, et al.
Veröffentlicht: (2024)
von: Song, Ruoyu, et al.
Veröffentlicht: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
von: Sharma, Mrinank, et al.
Veröffentlicht: (2025)
von: Sharma, Mrinank, et al.
Veröffentlicht: (2025)
Investigating the Impact of Dark Patterns on LLM-Based Web Agents
von: Ersoy, Devin, et al.
Veröffentlicht: (2025)
von: Ersoy, Devin, et al.
Veröffentlicht: (2025)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
von: Hasan, Adib, et al.
Veröffentlicht: (2024)
von: Hasan, Adib, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploring and Developing a Pre-Model Safeguard with Draft Models
von: Cai, Hongyu, et al.
Veröffentlicht: (2026) -
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
von: Hua, Peichun, et al.
Veröffentlicht: (2025) -
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024) -
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
von: Lin, Shi, et al.
Veröffentlicht: (2024) -
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)