On Jailbreaking Quantized Language Models Through Fault Injection Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zahran, Noureldin, Tahmasivand, Ahmad, Alouani, Ihsen, Khasawneh, Khaled, Fouda, Mohammed E. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models
von: Tahmasivand, Ahmad, et al.
Veröffentlicht: (2025)
von: Tahmasivand, Ahmad, et al.
Veröffentlicht: (2025)
Bypassing Prompt Injection Detectors through Evasive Injections
von: Rahman, Md Jahedur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Jahedur, et al.
Veröffentlicht: (2026)
Evasive Hardware Trojan through Adversarial Power Trace
von: Omidi, Behnam, et al.
Veröffentlicht: (2024)
von: Omidi, Behnam, et al.
Veröffentlicht: (2024)
SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models
von: Abdelnasser, Omar, et al.
Veröffentlicht: (2026)
von: Abdelnasser, Omar, et al.
Veröffentlicht: (2026)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Fault Injection and Safe-Error Attack for Extraction of Embedded Neural Network Models
von: Hector, Kevin, et al.
Veröffentlicht: (2023)
von: Hector, Kevin, et al.
Veröffentlicht: (2023)
Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update Analysis
von: Park, Jeonghwan, et al.
Veröffentlicht: (2025)
von: Park, Jeonghwan, et al.
Veröffentlicht: (2025)
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
von: Li, Jie, et al.
Veröffentlicht: (2024)
von: Li, Jie, et al.
Veröffentlicht: (2024)
SoK: Robustness in Large Language Models against Jailbreak Attacks
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
Untargeted Jailbreak Attack
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
von: Huang, Xinzhe, et al.
Veröffentlicht: (2025)
Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
von: Yang, Fan
Veröffentlicht: (2025)
von: Yang, Fan
Veröffentlicht: (2025)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
BrainLeaks: On the Privacy-Preserving Properties of Neuromorphic Architectures against Model Inversion Attacks
von: Poursiami, Hamed, et al.
Veröffentlicht: (2024)
von: Poursiami, Hamed, et al.
Veröffentlicht: (2024)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
von: Zhang, Wenhui, et al.
Veröffentlicht: (2025)
von: Zhang, Wenhui, et al.
Veröffentlicht: (2025)
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
von: Hong, Wenjing, et al.
Veröffentlicht: (2026)
von: Hong, Wenjing, et al.
Veröffentlicht: (2026)
PEFT-as-an-Attack! Jailbreaking Language Models during Federated Parameter-Efficient Fine-Tuning
von: Li, Shenghui, et al.
Veröffentlicht: (2024)
von: Li, Shenghui, et al.
Veröffentlicht: (2024)
AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks
von: Song, Weiming, et al.
Veröffentlicht: (2026)
von: Song, Weiming, et al.
Veröffentlicht: (2026)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
von: Teng, Ma, et al.
Veröffentlicht: (2024)
von: Teng, Ma, et al.
Veröffentlicht: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection
von: Wang, Zhilong, et al.
Veröffentlicht: (2024)
von: Wang, Zhilong, et al.
Veröffentlicht: (2024)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
von: Saha, Shoumik, et al.
Veröffentlicht: (2025)
von: Saha, Shoumik, et al.
Veröffentlicht: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
Jailbreaking Large Language Models through Iterative Tool-Disguised Attacks via Reinforcement Learning
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2026)
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2026)
Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models
von: Yang, Yiqi, et al.
Veröffentlicht: (2024)
von: Yang, Yiqi, et al.
Veröffentlicht: (2024)
GUARD-SLM: Token Activation-Based Defense Against Jailbreak Attacks for Small Language Models
von: Mia, Md Jueal, et al.
Veröffentlicht: (2026)
von: Mia, Md Jueal, et al.
Veröffentlicht: (2026)
Involuntary Jailbreak: On Self-Prompting Attacks
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
von: Guo, Yangyang, et al.
Veröffentlicht: (2025)
Distract Large Language Models for Automatic Jailbreak Attack
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
von: Zaree, Pedram, et al.
Veröffentlicht: (2025)
von: Zaree, Pedram, et al.
Veröffentlicht: (2025)
Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
von: Yan, Yu, et al.
Veröffentlicht: (2026)
von: Yan, Yu, et al.
Veröffentlicht: (2026)
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
von: Li, Zongze, et al.
Veröffentlicht: (2025)
von: Li, Zongze, et al.
Veröffentlicht: (2025)
FlipAttack: Jailbreak LLMs via Flipping
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
New Wide-Net-Casting Jailbreak Attacks Risk Large Models
von: Xiang, Qiuchi, et al.
Veröffentlicht: (2026)
von: Xiang, Qiuchi, et al.
Veröffentlicht: (2026)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2024)
von: Yu, Miao, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models Through Content Concretization
von: Wahréus, Johan, et al.
Veröffentlicht: (2025)
von: Wahréus, Johan, et al.
Veröffentlicht: (2025)
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
von: Ding, Renhua, et al.
Veröffentlicht: (2025)
von: Ding, Renhua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models
von: Tahmasivand, Ahmad, et al.
Veröffentlicht: (2025) -
Bypassing Prompt Injection Detectors through Evasive Injections
von: Rahman, Md Jahedur, et al.
Veröffentlicht: (2026) -
Evasive Hardware Trojan through Adversarial Power Trace
von: Omidi, Behnam, et al.
Veröffentlicht: (2024) -
SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models
von: Abdelnasser, Omar, et al.
Veröffentlicht: (2026) -
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
von: Park, Junyoung, et al.
Veröffentlicht: (2026)