Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Galinkin, Erick, Sablotny, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
von: Peng, Benji, et al.
Veröffentlicht: (2024)
von: Peng, Benji, et al.
Veröffentlicht: (2024)
Towards Type Agnostic Cyber Defense Agents
von: Galinkin, Erick, et al.
Veröffentlicht: (2024)
von: Galinkin, Erick, et al.
Veröffentlicht: (2024)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
von: Hua, Peichun, et al.
Veröffentlicht: (2025)
von: Hua, Peichun, et al.
Veröffentlicht: (2025)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
EnJa: Ensemble Jailbreak on Large Language Models
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2024)
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
The Price of Pessimism for Automated Defense
von: Galinkin, Erick, et al.
Veröffentlicht: (2024)
von: Galinkin, Erick, et al.
Veröffentlicht: (2024)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
BEACON: Behavioral Malware Classification with Large Language Model Embeddings and Deep Learning
von: Perera, Wadduwage Shanika, et al.
Veröffentlicht: (2025)
von: Perera, Wadduwage Shanika, et al.
Veröffentlicht: (2025)
Finetuning Large Language Models for Vulnerability Detection
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
von: Nikolić, Kristina, et al.
Veröffentlicht: (2025)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
von: Guo, Qiming, et al.
Veröffentlicht: (2025)
von: Guo, Qiming, et al.
Veröffentlicht: (2025)
Watermark Stealing in Large Language Models
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
von: Lin, Shi, et al.
Veröffentlicht: (2024)
von: Lin, Shi, et al.
Veröffentlicht: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
Rethinking How to Evaluate Language Model Jailbreak
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
Mitigating Many-Shot Jailbreaking
von: Ackerman, Christopher M., et al.
Veröffentlicht: (2025)
von: Ackerman, Christopher M., et al.
Veröffentlicht: (2025)
Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2023)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
von: Li, Xiao, et al.
Veröffentlicht: (2024)
von: Li, Xiao, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Phishing Detection, Self-Consistency, Faithfulness, and Explainability
von: Kuikel, Shova, et al.
Veröffentlicht: (2025)
von: Kuikel, Shova, et al.
Veröffentlicht: (2025)
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
von: Chen, Guangke, et al.
Veröffentlicht: (2025)
von: Chen, Guangke, et al.
Veröffentlicht: (2025)
An Explainable Transformer-based Model for Phishing Email Detection: A Large Language Model Approach
von: Uddin, Mohammad Amaz, et al.
Veröffentlicht: (2024)
von: Uddin, Mohammad Amaz, et al.
Veröffentlicht: (2024)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
von: Wang, Zi, et al.
Veröffentlicht: (2024)
von: Wang, Zi, et al.
Veröffentlicht: (2024)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
Reconstruction of Differentially Private Text Sanitization via Large Language Models
von: Pang, Shuchao, et al.
Veröffentlicht: (2024)
von: Pang, Shuchao, et al.
Veröffentlicht: (2024)
Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models
von: Wang, Jeffrey G., et al.
Veröffentlicht: (2024)
von: Wang, Jeffrey G., et al.
Veröffentlicht: (2024)
Advancing Email Spam Detection: Leveraging Zero-Shot Learning and Large Language Models
von: SHirvani, Ghazaleh, et al.
Veröffentlicht: (2025)
von: SHirvani, Ghazaleh, et al.
Veröffentlicht: (2025)
Security and Detectability Analysis of Unicode Text Watermarking Methods Against Large Language Models
von: Hellmeier, Malte
Veröffentlicht: (2025)
von: Hellmeier, Malte
Veröffentlicht: (2025)
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
von: Goel, Aman, et al.
Veröffentlicht: (2025)
von: Goel, Aman, et al.
Veröffentlicht: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models
von: Krishna, Arjun, et al.
Veröffentlicht: (2025) -
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
von: Peng, Benji, et al.
Veröffentlicht: (2024) -
Towards Type Agnostic Cyber Defense Agents
von: Galinkin, Erick, et al.
Veröffentlicht: (2024) -
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
von: Hua, Peichun, et al.
Veröffentlicht: (2025) -
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)