Saved in:
| Main Authors: | Jiralerspong, Thomas, Kondrup, Flemming, Bengio, Yoshua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.16928 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025)
by: Williams-King, David, et al.
Published: (2025)
Visual CoT Makes VLMs Smarter but More Fragile
by: Xu, Chunxue, et al.
Published: (2025)
by: Xu, Chunxue, et al.
Published: (2025)
AgentWatcher: A Rule-based Prompt Injection Monitor
by: Wang, Yanting, et al.
Published: (2026)
by: Wang, Yanting, et al.
Published: (2026)
Can We Infer Confidential Properties of Training Data from LLMs?
by: Huang, Pengrun, et al.
Published: (2025)
by: Huang, Pengrun, et al.
Published: (2025)
Reliable Weak-to-Strong Monitoring of LLM Agents
by: Kale, Neil, et al.
Published: (2025)
by: Kale, Neil, et al.
Published: (2025)
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
by: He, Pengfei, et al.
Published: (2026)
by: He, Pengfei, et al.
Published: (2026)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Inferring Sensitive Attributes from Knowledge Graph Embeddings: Attack and Defense Strategies
by: Hayder, Yasmine
Published: (2026)
by: Hayder, Yasmine
Published: (2026)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
BlockDoor: Blocking Backdoor Based Watermarks in Deep Neural Networks
by: Puah, Yi Hao, et al.
Published: (2024)
by: Puah, Yi Hao, et al.
Published: (2024)
Memory-Induced Tool-Drift in LLM Agents
by: Dabas, Mahavir, et al.
Published: (2026)
by: Dabas, Mahavir, et al.
Published: (2026)
Improving LLM-Assisted Secure Code Generation through Retrieval-Augmented-Generation and Multi-Tool Feedback
by: Sriram, Vidyut, et al.
Published: (2026)
by: Sriram, Vidyut, et al.
Published: (2026)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
by: Lee, Dongjun, et al.
Published: (2026)
by: Lee, Dongjun, et al.
Published: (2026)
Design Patterns for Securing LLM Agents against Prompt Injections
by: Beurer-Kellner, Luca, et al.
Published: (2025)
by: Beurer-Kellner, Luca, et al.
Published: (2025)
CoT-Guard: Small Models for Strong Monitoring
by: Diwan, Nirav, et al.
Published: (2026)
by: Diwan, Nirav, et al.
Published: (2026)
Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
by: Ahmed, Nesreen K., et al.
Published: (2026)
by: Ahmed, Nesreen K., et al.
Published: (2026)
SunBlock: Cloudless Protection for IoT Systems
by: Safronov, Vadim, et al.
Published: (2024)
by: Safronov, Vadim, et al.
Published: (2024)
Rerouting LLM Routers
by: Shafran, Avital, et al.
Published: (2025)
by: Shafran, Avital, et al.
Published: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
by: Hossain, S M Asif, et al.
Published: (2025)
by: Hossain, S M Asif, et al.
Published: (2025)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
by: Yamamoto, Yuji, et al.
Published: (2026)
by: Yamamoto, Yuji, et al.
Published: (2026)
Denoising the US Census: Succinct Block Hierarchical Regression
by: Ghazi, Badih, et al.
Published: (2026)
by: Ghazi, Badih, et al.
Published: (2026)
Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents
by: Tran, Toan, et al.
Published: (2026)
by: Tran, Toan, et al.
Published: (2026)
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
by: Lee, Hwiwon, et al.
Published: (2025)
by: Lee, Hwiwon, et al.
Published: (2025)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
by: Zhan, Qiusi, et al.
Published: (2025)
by: Zhan, Qiusi, et al.
Published: (2025)
On the Generalizability of Machine Learning-based Ransomware Detection in Block Storage
by: Reategui, Nicolas, et al.
Published: (2024)
by: Reategui, Nicolas, et al.
Published: (2024)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
CoTDeceptor:Adversarial Code Obfuscation Against CoT-Enhanced LLM Code Agents
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
DP-Dueling: Learning from Preference Feedback without Compromising User Privacy
by: Saha, Aadirupa, et al.
Published: (2024)
by: Saha, Aadirupa, et al.
Published: (2024)
Your Agent Can Defend Itself against Backdoor Attacks
by: Changjiang, Li, et al.
Published: (2025)
by: Changjiang, Li, et al.
Published: (2025)
Can Copyright be Reduced to Privacy?
by: Elkin-Koren, Niva, et al.
Published: (2023)
by: Elkin-Koren, Niva, et al.
Published: (2023)
XChainWatcher: Monitoring and Identifying Attacks in Cross-Chain Bridges
by: Augusto, André, et al.
Published: (2024)
by: Augusto, André, et al.
Published: (2024)
Cascade: Token-Sharded Private LLM Inference
by: Thomas, Rahul, et al.
Published: (2025)
by: Thomas, Rahul, et al.
Published: (2025)
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
by: Chu, Kexin
Published: (2026)
by: Chu, Kexin
Published: (2026)
Can LLMs Patch Security Issues?
by: Alrashedy, Kamel, et al.
Published: (2023)
by: Alrashedy, Kamel, et al.
Published: (2023)
A Machine Learning-Based Framework for Assessing Cryptographic Indistinguishability of Lightweight Block Ciphers
by: Dani, Jimmy, et al.
Published: (2024)
by: Dani, Jimmy, et al.
Published: (2024)
Differentially Private SGD Without Clipping Bias: An Error-Feedback Approach
by: Zhang, Xinwei, et al.
Published: (2023)
by: Zhang, Xinwei, et al.
Published: (2023)
Scalable Neural Network Verification with Branch-and-bound Inferred Cutting Planes
by: Zhou, Duo, et al.
Published: (2024)
by: Zhou, Duo, et al.
Published: (2024)
Similar Items
-
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025) -
Visual CoT Makes VLMs Smarter but More Fragile
by: Xu, Chunxue, et al.
Published: (2025) -
AgentWatcher: A Rule-based Prompt Injection Monitor
by: Wang, Yanting, et al.
Published: (2026) -
Can We Infer Confidential Properties of Training Data from LLMs?
by: Huang, Pengrun, et al.
Published: (2025) -
Reliable Weak-to-Strong Monitoring of LLM Agents
by: Kale, Neil, et al.
Published: (2025)