Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Haoyu, Zandsalimy, Mohammad, Sushmita, Shanu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models
von: Zheng, Youjia, et al.
Veröffentlicht: (2025)
von: Zheng, Youjia, et al.
Veröffentlicht: (2025)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
von: Vega, Jason, et al.
Veröffentlicht: (2023)
von: Vega, Jason, et al.
Veröffentlicht: (2023)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
Lifelong Safety Alignment for Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
von: Wang, Kun, et al.
Veröffentlicht: (2025)
von: Wang, Kun, et al.
Veröffentlicht: (2025)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
von: Shetty, Anudeex, et al.
Veröffentlicht: (2026)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2026)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
von: Jin, Chang, et al.
Veröffentlicht: (2026)
von: Jin, Chang, et al.
Veröffentlicht: (2026)
Kov: Transferable and Naturalistic Black-Box LLM Attacks using Markov Decision Processes and Tree Search
von: Moss, Robert J.
Veröffentlicht: (2024)
von: Moss, Robert J.
Veröffentlicht: (2024)
A Semantic and Clean-label Backdoor Attack against Graph Convolutional Networks
von: Dai, Jiazhu, et al.
Veröffentlicht: (2025)
von: Dai, Jiazhu, et al.
Veröffentlicht: (2025)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
von: Li, Xueyi, et al.
Veröffentlicht: (2026)
von: Li, Xueyi, et al.
Veröffentlicht: (2026)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
von: Correia, Pedro H. Barcha, et al.
Veröffentlicht: (2026)
von: Correia, Pedro H. Barcha, et al.
Veröffentlicht: (2026)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
Mind the Gap: A Practical Attack on GGUF Quantization
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
von: Wei, Zeming, et al.
Veröffentlicht: (2026)
von: Wei, Zeming, et al.
Veröffentlicht: (2026)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Membership Inference Attack with Partial Features
von: Wang, Xurun, et al.
Veröffentlicht: (2025)
von: Wang, Xurun, et al.
Veröffentlicht: (2025)
On the Role of Attention Heads in Large Language Model Safety
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
Membership Inference Attacks on LLM-based Recommender Systems
von: He, Jiajie, et al.
Veröffentlicht: (2025)
von: He, Jiajie, et al.
Veröffentlicht: (2025)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
von: Lu, Ning, et al.
Veröffentlicht: (2023)
von: Lu, Ning, et al.
Veröffentlicht: (2023)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
von: Srivastava, Saksham Sahai, et al.
Veröffentlicht: (2025)
von: Srivastava, Saksham Sahai, et al.
Veröffentlicht: (2025)
Intent Laundering: AI Safety Datasets Are Not What They Seem
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models
von: Zheng, Youjia, et al.
Veröffentlicht: (2025) -
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2026) -
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
von: Vega, Jason, et al.
Veröffentlicht: (2023) -
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024) -
Lifelong Safety Alignment for Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)