Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Haochuan Kevin, Zhang, Zechen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Whisper Leak: a side-channel attack on Large Language Models
by: McDonald, Geoff, et al.
Published: (2025)
by: McDonald, Geoff, et al.
Published: (2025)
Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
by: Dar, Daniyal Kabir, et al.
Published: (2025)
by: Dar, Daniyal Kabir, et al.
Published: (2025)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
by: Doda, Shravan
Published: (2026)
by: Doda, Shravan
Published: (2026)
Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models
by: Ntais, Pavlos
Published: (2025)
by: Ntais, Pavlos
Published: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026)
by: Othman, Refat
Published: (2026)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
by: Gu, Yongtong, et al.
Published: (2026)
by: Gu, Yongtong, et al.
Published: (2026)
VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use
by: Santillana, Juan S.
Published: (2026)
by: Santillana, Juan S.
Published: (2026)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
by: Morales, Jaime, et al.
Published: (2026)
by: Morales, Jaime, et al.
Published: (2026)
Can Graph-Based Microservice Performance Detection Be Used for Microservice Intrusion Detection?
by: Ma, Yunjian
Published: (2026)
by: Ma, Yunjian
Published: (2026)
Phishing Detection System: An Ensemble Approach Using Character-Level CNN and Feature Engineering
by: Dubey, Rudra, et al.
Published: (2025)
by: Dubey, Rudra, et al.
Published: (2025)
h4rm3l: A language for Composable Jailbreak Attack Synthesis
by: Doumbouya, Moussa Koulako Bala, et al.
Published: (2024)
by: Doumbouya, Moussa Koulako Bala, et al.
Published: (2024)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
by: Zanbaghi, Shahin, et al.
Published: (2025)
by: Zanbaghi, Shahin, et al.
Published: (2025)
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought
by: Zhang, Yuyi, et al.
Published: (2025)
by: Zhang, Yuyi, et al.
Published: (2025)
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
by: Qi, Jinhu, et al.
Published: (2026)
by: Qi, Jinhu, et al.
Published: (2026)
A Protocol-Language Model for Network Intrusion (Without Deep Packet Inspection)
by: Sharma, Vivek Kumar
Published: (2026)
by: Sharma, Vivek Kumar
Published: (2026)
Generalizable and Interpretable RF Fingerprinting with Shapelet-Enhanced Large Language Models
by: Zhao, Tianya, et al.
Published: (2026)
by: Zhao, Tianya, et al.
Published: (2026)
CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
VoiceSHIELD-Small: Real-Time Malicious Speech Detection and Transcription
by: Ranjan, Sumit, et al.
Published: (2026)
by: Ranjan, Sumit, et al.
Published: (2026)
Towards Modeling Cybersecurity Behavior of Humans in Organizations
by: Kürtz, Klaas Ole
Published: (2026)
by: Kürtz, Klaas Ole
Published: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025)
by: Dawson, Ads, et al.
Published: (2025)
The Automation Advantage in AI Red Teaming
by: Mulla, Rob, et al.
Published: (2025)
by: Mulla, Rob, et al.
Published: (2025)
Robustness, Cost, and Attack-Surface Concentration in Phishing Detection
by: Allagan, Julian, et al.
Published: (2026)
by: Allagan, Julian, et al.
Published: (2026)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
Send to which account? Evaluation of an LLM-based Scambaiting System
by: Siadati, Hossein, et al.
Published: (2025)
by: Siadati, Hossein, et al.
Published: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Measuring Harmfulness of Computer-Using Agents
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
Countermind: A Multi-Layered Security Architecture for Large Language Models
by: Schwarz, Dominik
Published: (2025)
by: Schwarz, Dominik
Published: (2025)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values
by: Chun, Jon, et al.
Published: (2024)
by: Chun, Jon, et al.
Published: (2024)
Robust Uncertainty Quantification for Factual Generation of Large Language Models
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
Lightweight LLMs for Network Attack Detection in IoT Networks
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026)
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026)
Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
by: Adjei, Yaw Osei, et al.
Published: (2025)
by: Adjei, Yaw Osei, et al.
Published: (2025)
Exploiting Web Search Tools of AI Agents for Data Exfiltration
by: Rall, Dennis, et al.
Published: (2025)
by: Rall, Dennis, et al.
Published: (2025)
Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
by: Subedar, Noah, et al.
Published: (2025)
by: Subedar, Noah, et al.
Published: (2025)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
DeRAG: Black-box Adversarial Attacks on Multiple Retrieval-Augmented Generation Applications via Prompt Injection
by: Wang, Jerry, et al.
Published: (2025)
by: Wang, Jerry, et al.
Published: (2025)
Similar Items
-
Whisper Leak: a side-channel attack on Large Language Models
by: McDonald, Geoff, et al.
Published: (2025) -
Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
by: Dar, Daniyal Kabir, et al.
Published: (2025) -
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
by: Doda, Shravan
Published: (2026) -
Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models
by: Ntais, Pavlos
Published: (2025) -
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026)