Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kumar, Shubham, Ahuja, Narendra |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Measuring the (Un)Faithfulness of Concept-Based Explanations
par: Kumar, Shubham, et autres
Publié: (2025)
par: Kumar, Shubham, et autres
Publié: (2025)
LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
par: Bhatnagar, Shubhang, et autres
Publié: (2025)
par: Bhatnagar, Shubhang, et autres
Publié: (2025)
Locally-Minimal Probabilistic Explanations
par: Izza, Yacine, et autres
Publié: (2023)
par: Izza, Yacine, et autres
Publié: (2023)
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
par: Ramesh, Govind, et autres
Publié: (2024)
par: Ramesh, Govind, et autres
Publié: (2024)
Causality-Aware Local Interpretable Model-Agnostic Explanations
par: Cinquini, Martina, et autres
Publié: (2022)
par: Cinquini, Martina, et autres
Publié: (2022)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
par: Ball, Sarah, et autres
Publié: (2024)
par: Ball, Sarah, et autres
Publié: (2024)
Piecewise-Linear Manifolds for Deep Metric Learning
par: Bhatnagar, Shubhang, et autres
Publié: (2024)
par: Bhatnagar, Shubhang, et autres
Publié: (2024)
Boosting Jailbreak Transferability for Large Language Models
par: Liu, Hanqing, et autres
Publié: (2024)
par: Liu, Hanqing, et autres
Publié: (2024)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
par: Zhou, Weikang, et autres
Publié: (2024)
par: Zhou, Weikang, et autres
Publié: (2024)
Potential Field Based Deep Metric Learning
par: Bhatnagar, Shubhang, et autres
Publié: (2024)
par: Bhatnagar, Shubhang, et autres
Publié: (2024)
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
par: Sharma, Shubham, et autres
Publié: (2025)
par: Sharma, Shubham, et autres
Publié: (2025)
Jailbreaking Large Language Models with Symbolic Mathematics
par: Bethany, Emet, et autres
Publié: (2024)
par: Bethany, Emet, et autres
Publié: (2024)
Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
par: Maple, Carsten, et autres
Publié: (2026)
par: Maple, Carsten, et autres
Publié: (2026)
Emoji-Based Jailbreaking of Large Language Models
par: Gopinadh, M P V S, et autres
Publié: (2026)
par: Gopinadh, M P V S, et autres
Publié: (2026)
Jailbreaking Large Vision Language Models in Intelligent Transportation Systems
par: Das, Badhan Chandra, et autres
Publié: (2025)
par: Das, Badhan Chandra, et autres
Publié: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
par: Tu, Shangqing, et autres
Publié: (2024)
par: Tu, Shangqing, et autres
Publié: (2024)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
par: Lee, Isack, et autres
Publié: (2024)
par: Lee, Isack, et autres
Publié: (2024)
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
par: Song, Zirui, et autres
Publié: (2025)
par: Song, Zirui, et autres
Publié: (2025)
Improving Tool Retrieval by Leveraging Large Language Models for Query Generation
par: Kachuee, Mohammad, et autres
Publié: (2024)
par: Kachuee, Mohammad, et autres
Publié: (2024)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
par: Peng, Benji, et autres
Publié: (2024)
par: Peng, Benji, et autres
Publié: (2024)
Imperceptible Jailbreaking against Large Language Models
par: Gao, Kuofeng, et autres
Publié: (2025)
par: Gao, Kuofeng, et autres
Publié: (2025)
Large Causal Models from Large Language Models
par: Mahadevan, Sridhar
Publié: (2025)
par: Mahadevan, Sridhar
Publié: (2025)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
par: Li, Jie, et autres
Publié: (2024)
par: Li, Jie, et autres
Publié: (2024)
Rethinking Prompting Strategies for Multi-Label Recognition with Partial Annotations
par: Rawlekar, Samyak, et autres
Publié: (2024)
par: Rawlekar, Samyak, et autres
Publié: (2024)
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
par: Yu, Jiahao, et autres
Publié: (2023)
par: Yu, Jiahao, et autres
Publié: (2023)
Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models
par: Wit, Maria Carolina Cornelia, et autres
Publié: (2025)
par: Wit, Maria Carolina Cornelia, et autres
Publié: (2025)
ShallowJail: Steering Jailbreaks against Large Language Models
par: Liu, Shang, et autres
Publié: (2026)
par: Liu, Shang, et autres
Publié: (2026)
The Cost of Thinking: Increased Jailbreak Risk in Large Language Models
par: Yang, Fan
Publié: (2025)
par: Yang, Fan
Publié: (2025)
Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models
par: Zheng, Youjia, et autres
Publié: (2025)
par: Zheng, Youjia, et autres
Publié: (2025)
Jailbreaking Black Box Large Language Models in Twenty Queries
par: Chao, Patrick, et autres
Publié: (2023)
par: Chao, Patrick, et autres
Publié: (2023)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
par: Wang, Xunguang, et autres
Publié: (2025)
par: Wang, Xunguang, et autres
Publié: (2025)
Causal Explanations for Image Classifiers
par: Chockler, Hana, et autres
Publié: (2024)
par: Chockler, Hana, et autres
Publié: (2024)
A Framework for Causal Concept-based Model Explanations
par: Bjøru, Anna Rodum, et autres
Publié: (2025)
par: Bjøru, Anna Rodum, et autres
Publié: (2025)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
par: Ran, Delong, et autres
Publié: (2024)
par: Ran, Delong, et autres
Publié: (2024)
Large Language Models as Nondeterministic Causal Models
par: Beckers, Sander
Publié: (2025)
par: Beckers, Sander
Publié: (2025)
Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
par: Xu, Xu, et autres
Publié: (2025)
par: Xu, Xu, et autres
Publié: (2025)
Jailbreaking Large Language Models Through Content Concretization
par: Wahréus, Johan, et autres
Publié: (2025)
par: Wahréus, Johan, et autres
Publié: (2025)
Distract Large Language Models for Automatic Jailbreak Attack
par: Xiao, Zeguan, et autres
Publié: (2024)
par: Xiao, Zeguan, et autres
Publié: (2024)
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
par: Zhao, Wei, et autres
Publié: (2024)
par: Zhao, Wei, et autres
Publié: (2024)
SoK: Robustness in Large Language Models against Jailbreak Attacks
par: Xu, Feiyue, et autres
Publié: (2026)
par: Xu, Feiyue, et autres
Publié: (2026)
Documents similaires
-
Measuring the (Un)Faithfulness of Concept-Based Explanations
par: Kumar, Shubham, et autres
Publié: (2025) -
LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
par: Bhatnagar, Shubhang, et autres
Publié: (2025) -
Locally-Minimal Probabilistic Explanations
par: Izza, Yacine, et autres
Publié: (2023) -
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
par: Ramesh, Govind, et autres
Publié: (2024) -
Causality-Aware Local Interpretable Model-Agnostic Explanations
par: Cinquini, Martina, et autres
Publié: (2022)