Single-pass Detection of Jailbreaking Input in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Candogan, Leyla Naz, Wu, Yongtao, Rocamora, Elias Abad, Chrysos, Grigorios G., Cevher, Volkan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Revisiting Character-level Adversarial Attacks for Language Models
par: Rocamora, Elias Abad, et autres
Publié: (2024)
par: Rocamora, Elias Abad, et autres
Publié: (2024)
Certified Robustness Under Bounded Levenshtein Distance
par: Rocamora, Elias Abad, et autres
Publié: (2025)
par: Rocamora, Elias Abad, et autres
Publié: (2025)
Linear Attention for Efficient Bidirectional Sequence Modeling
par: Afzal, Arshia, et autres
Publié: (2025)
par: Afzal, Arshia, et autres
Publié: (2025)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
par: Cheng, Yixin, et autres
Publié: (2024)
par: Cheng, Yixin, et autres
Publié: (2024)
Membership Inference Attacks against Large Vision-Language Models
par: Li, Zhan, et autres
Publié: (2024)
par: Li, Zhan, et autres
Publié: (2024)
Efficient local linearity regularization to overcome catastrophic overfitting
par: Rocamora, Elias Abad, et autres
Publié: (2024)
par: Rocamora, Elias Abad, et autres
Publié: (2024)
Robust NAS under adversarial training: benchmark, theory, and beyond
par: Wu, Yongtao, et autres
Publié: (2024)
par: Wu, Yongtao, et autres
Publié: (2024)
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
par: Afzal, Arshia, et autres
Publié: (2024)
par: Afzal, Arshia, et autres
Publié: (2024)
Robustness in Both Domains: CLIP Needs a Robust Text Encoder
par: Rocamora, Elias Abad, et autres
Publié: (2025)
par: Rocamora, Elias Abad, et autres
Publié: (2025)
Quantum-PEFT: Ultra parameter-efficient fine-tuning
par: Koike-Akino, Toshiaki, et autres
Publié: (2025)
par: Koike-Akino, Toshiaki, et autres
Publié: (2025)
Going beyond Compositions, DDPMs Can Produce Zero-Shot Interpolations
par: Deschenaux, Justin, et autres
Publié: (2024)
par: Deschenaux, Justin, et autres
Publié: (2024)
Hadamard product in deep learning: Introduction, Advances and Challenges
par: Chrysos, Grigorios G, et autres
Publié: (2025)
par: Chrysos, Grigorios G, et autres
Publié: (2025)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
par: Wu, Yongtao, et autres
Publié: (2025)
par: Wu, Yongtao, et autres
Publié: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
par: Erdogan, Mete, et autres
Publié: (2025)
par: Erdogan, Mete, et autres
Publié: (2025)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
par: Lee, Isack, et autres
Publié: (2024)
par: Lee, Isack, et autres
Publié: (2024)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
par: Afzal, Arshia, et autres
Publié: (2025)
par: Afzal, Arshia, et autres
Publié: (2025)
Multilinear Operator Networks
par: Cheng, Yixin, et autres
Publié: (2024)
par: Cheng, Yixin, et autres
Publié: (2024)
Learning to Remove Cuts in Integer Linear Programming
par: Puigdemont, Pol, et autres
Publié: (2024)
par: Puigdemont, Pol, et autres
Publié: (2024)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
par: Hua, Peichun, et autres
Publié: (2025)
par: Hua, Peichun, et autres
Publié: (2025)
The Last Mile to Supervised Performance: Semi-Supervised Domain Adaptation for Semantic Segmentation
par: Morales-Brotons, Daniel, et autres
Publié: (2024)
par: Morales-Brotons, Daniel, et autres
Publié: (2024)
Jailbreaking Large Language Models with Symbolic Mathematics
par: Bethany, Emet, et autres
Publié: (2024)
par: Bethany, Emet, et autres
Publié: (2024)
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
par: Wang, Xunguang, et autres
Publié: (2025)
par: Wang, Xunguang, et autres
Publié: (2025)
GRASP: Deterministic argument ranking in interaction graphs
par: Misra, Diganta, et autres
Publié: (2026)
par: Misra, Diganta, et autres
Publié: (2026)
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
par: Mura, Raffaele, et autres
Publié: (2025)
par: Mura, Raffaele, et autres
Publié: (2025)
EnJa: Ensemble Jailbreak on Large Language Models
par: Zhang, Jiahao, et autres
Publié: (2024)
par: Zhang, Jiahao, et autres
Publié: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
par: Hu, Xiaomeng, et autres
Publié: (2024)
par: Hu, Xiaomeng, et autres
Publié: (2024)
Generalization of Scaled Deep ResNets in the Mean-Field Regime
par: Chen, Yihang, et autres
Publié: (2024)
par: Chen, Yihang, et autres
Publié: (2024)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
par: Ball, Sarah, et autres
Publié: (2024)
par: Ball, Sarah, et autres
Publié: (2024)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
par: Xu, Zihao, et autres
Publié: (2024)
par: Xu, Zihao, et autres
Publié: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
par: Oh, Sejoon, et autres
Publié: (2024)
par: Oh, Sejoon, et autres
Publié: (2024)
Talking Nonsense: Probing Large Language Models' Understanding of Adversarial Gibberish Inputs
par: Cherepanova, Valeriia, et autres
Publié: (2024)
par: Cherepanova, Valeriia, et autres
Publié: (2024)
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination
par: Song, Jongyoon, et autres
Publié: (2024)
par: Song, Jongyoon, et autres
Publié: (2024)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
par: Boreiko, Valentyn, et autres
Publié: (2024)
par: Boreiko, Valentyn, et autres
Publié: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
par: Yi, Sibo, et autres
Publié: (2024)
par: Yi, Sibo, et autres
Publié: (2024)
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
par: Goel, Aman, et autres
Publié: (2025)
par: Goel, Aman, et autres
Publié: (2025)
Adversarial Training for Defense Against Label Poisoning Attacks
par: Bal, Melis Ilayda, et autres
Publié: (2025)
par: Bal, Melis Ilayda, et autres
Publié: (2025)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
par: Zhou, Weikang, et autres
Publié: (2024)
par: Zhou, Weikang, et autres
Publié: (2024)
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
par: Bisconti, Piercosma, et autres
Publié: (2025)
par: Bisconti, Piercosma, et autres
Publié: (2025)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
par: Lin, Shi, et autres
Publié: (2024)
par: Lin, Shi, et autres
Publié: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
par: Reddy, Aashray, et autres
Publié: (2025)
par: Reddy, Aashray, et autres
Publié: (2025)
Documents similaires
-
Revisiting Character-level Adversarial Attacks for Language Models
par: Rocamora, Elias Abad, et autres
Publié: (2024) -
Certified Robustness Under Bounded Levenshtein Distance
par: Rocamora, Elias Abad, et autres
Publié: (2025) -
Linear Attention for Efficient Bidirectional Sequence Modeling
par: Afzal, Arshia, et autres
Publié: (2025) -
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
par: Cheng, Yixin, et autres
Publié: (2024) -
Membership Inference Attacks against Large Vision-Language Models
par: Li, Zhan, et autres
Publié: (2024)