GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Sinan, Wang, An |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs
von: Fei, Zekun, et al.
Veröffentlicht: (2026)
von: Fei, Zekun, et al.
Veröffentlicht: (2026)
Rethinking Fraud Safety Evaluation: Multi-Round Attacks Reveal Safety-Utility Tradeoffs in Graph-Context LLM Defenders
von: Jiang, Laura, et al.
Veröffentlicht: (2026)
von: Jiang, Laura, et al.
Veröffentlicht: (2026)
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
von: Li, Junchen, et al.
Veröffentlicht: (2026)
von: Li, Junchen, et al.
Veröffentlicht: (2026)
Quantization Blindspots: How Model Compression Breaks Backdoor Defenses
von: Pandey, Rohan, et al.
Veröffentlicht: (2025)
von: Pandey, Rohan, et al.
Veröffentlicht: (2025)
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
von: Lian, Zhuotao, et al.
Veröffentlicht: (2025)
von: Lian, Zhuotao, et al.
Veröffentlicht: (2025)
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
von: Liu, Zesen, et al.
Veröffentlicht: (2025)
von: Liu, Zesen, et al.
Veröffentlicht: (2025)
Exploiting Leakage in Password Managers via Injection Attacks
von: Fábrega, Andrés, et al.
Veröffentlicht: (2024)
von: Fábrega, Andrés, et al.
Veröffentlicht: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
Physical-Layer Signal Injection Attacks on EV Charging Ports: Bypassing Authentication via Electrical-Level Exploits
von: Shi, Hetian, et al.
Veröffentlicht: (2025)
von: Shi, Hetian, et al.
Veröffentlicht: (2025)
CAN-Trace Attack: Exploit CAN Messages to Uncover Driving Trajectories
von: Lin, Xiaojie, et al.
Veröffentlicht: (2025)
von: Lin, Xiaojie, et al.
Veröffentlicht: (2025)
Exploiting AI for Attacks: On the Interplay between Adversarial AI and Offensive AI
von: Schröer, Saskia Laura, et al.
Veröffentlicht: (2025)
von: Schröer, Saskia Laura, et al.
Veröffentlicht: (2025)
Exploring and Exploiting the Resource Isolation Attack Surface of WebAssembly Containers
von: Yu, Zhaofeng, et al.
Veröffentlicht: (2025)
von: Yu, Zhaofeng, et al.
Veröffentlicht: (2025)
Exploiting Liquidity Exhaustion Attacks in Intent-Based Cross-Chain Bridges
von: Augusto, André, et al.
Veröffentlicht: (2026)
von: Augusto, André, et al.
Veröffentlicht: (2026)
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
von: Zhang, Junke, et al.
Veröffentlicht: (2026)
von: Zhang, Junke, et al.
Veröffentlicht: (2026)
FLARE: Fault Attack Leveraging Address Reconfiguration Exploits in Multi-Tenant FPGAs
von: Chaudhuri, Jayeeta, et al.
Veröffentlicht: (2025)
von: Chaudhuri, Jayeeta, et al.
Veröffentlicht: (2025)
MultiKG: Multi-Source Threat Intelligence Aggregation for High-Quality Knowledge Graph Representation of Attack Techniques
von: Wang, Jian, et al.
Veröffentlicht: (2024)
von: Wang, Jian, et al.
Veröffentlicht: (2024)
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Exploiting LLM Quantization
von: Egashira, Kazuki, et al.
Veröffentlicht: (2024)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2024)
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
von: Zhang, Qingzhao, et al.
Veröffentlicht: (2024)
von: Zhang, Qingzhao, et al.
Veröffentlicht: (2024)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
von: He, Zeqing, et al.
Veröffentlicht: (2024)
von: He, Zeqing, et al.
Veröffentlicht: (2024)
When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
von: Li, Jiaqing, et al.
Veröffentlicht: (2026)
von: Li, Jiaqing, et al.
Veröffentlicht: (2026)
GATEBLEED: Exploiting On-Core Accelerator Power Gating for High Performance & Stealthy Attacks on AI
von: Kalyanapu, Joshua, et al.
Veröffentlicht: (2025)
von: Kalyanapu, Joshua, et al.
Veröffentlicht: (2025)
SleepWalk: Exploiting Context Switching and Residual Power for Physical Side-Channel Attacks
von: Sanjaya, Sahan, et al.
Veröffentlicht: (2025)
von: Sanjaya, Sahan, et al.
Veröffentlicht: (2025)
Exploiting Cross-Layer Vulnerabilities: Off-Path Attacks on the TCP/IP Protocol Suite
von: Feng, Xuewei, et al.
Veröffentlicht: (2024)
von: Feng, Xuewei, et al.
Veröffentlicht: (2024)
QUIC-Exfil: Exploiting QUIC's Server Preferred Address Feature to Perform Data Exfiltration Attacks
von: Grübl, Thomas, et al.
Veröffentlicht: (2025)
von: Grübl, Thomas, et al.
Veröffentlicht: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
von: Yin, Yu, et al.
Veröffentlicht: (2026)
von: Yin, Yu, et al.
Veröffentlicht: (2026)
Exploiting Inaccurate Branch History in Side-Channel Attacks
von: Zhu, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhu, Yuhui, et al.
Veröffentlicht: (2025)
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
von: He, Pengfei, et al.
Veröffentlicht: (2025)
von: He, Pengfei, et al.
Veröffentlicht: (2025)
Generalized Adversarial Code-Suggestions: Exploiting Contexts of LLM-based Code-Completion
von: Rubel, Karl, et al.
Veröffentlicht: (2024)
von: Rubel, Karl, et al.
Veröffentlicht: (2024)
Provably Robust Explainable Graph Neural Networks against Graph Perturbation Attacks
von: Li, Jiate, et al.
Veröffentlicht: (2025)
von: Li, Jiate, et al.
Veröffentlicht: (2025)
Commitment Attacks on Ethereum's Reward Mechanism
von: Sarenche, Roozbeh, et al.
Veröffentlicht: (2024)
von: Sarenche, Roozbeh, et al.
Veröffentlicht: (2024)
Attributing and Exploiting Safety Vectors through Global Optimization in Large Language Models
von: Chu, Fengheng, et al.
Veröffentlicht: (2026)
von: Chu, Fengheng, et al.
Veröffentlicht: (2026)
Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis
von: Topol, Zvi
Veröffentlicht: (2026)
von: Topol, Zvi
Veröffentlicht: (2026)
Finding Software Supply Chain Attack Paths with Logical Attack Graphs
von: Soeiro, Luıs, et al.
Veröffentlicht: (2025)
von: Soeiro, Luıs, et al.
Veröffentlicht: (2025)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
von: An, Li, et al.
Veröffentlicht: (2025)
von: An, Li, et al.
Veröffentlicht: (2025)
Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4
von: Polyakov, Alex, et al.
Veröffentlicht: (2026)
von: Polyakov, Alex, et al.
Veröffentlicht: (2026)
Cluster-Aware Attacks on Graph Watermarks
von: Nemecek, Alexander, et al.
Veröffentlicht: (2025)
von: Nemecek, Alexander, et al.
Veröffentlicht: (2025)
Data Poisoning Attacks to Local Differential Privacy Protocols for Graphs
von: He, Xi, et al.
Veröffentlicht: (2024)
von: He, Xi, et al.
Veröffentlicht: (2024)
System Password Security: Attack and Defense Mechanisms
von: Shi, Chaofang, et al.
Veröffentlicht: (2025)
von: Shi, Chaofang, et al.
Veröffentlicht: (2025)
What Hard Tokens Reveal: Exploiting Low-confidence Tokens for Membership Inference Attacks against Large Language Models
von: Jawad, Md Tasnim, et al.
Veröffentlicht: (2026)
von: Jawad, Md Tasnim, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs
von: Fei, Zekun, et al.
Veröffentlicht: (2026) -
Rethinking Fraud Safety Evaluation: Multi-Round Attacks Reveal Safety-Utility Tradeoffs in Graph-Context LLM Defenders
von: Jiang, Laura, et al.
Veröffentlicht: (2026) -
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
von: Li, Junchen, et al.
Veröffentlicht: (2026) -
Quantization Blindspots: How Model Compression Breaks Backdoor Defenses
von: Pandey, Rohan, et al.
Veröffentlicht: (2025) -
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
von: Lian, Zhuotao, et al.
Veröffentlicht: (2025)