Breach By A Thousand Leaks: Unsafe Information Leakage in `Safe' AI Responses
Fuente:
arXiv
Saved in:
| Main Authors: | Glukhov, David, Han, Ziwen, Shumailov, Ilia, Papyan, Vardan, Papernot, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
by: Thudi, Anvith, et al.
Published: (2023)
by: Thudi, Anvith, et al.
Published: (2023)
Architectural Neural Backdoors from First Principles
by: Langford, Harry, et al.
Published: (2024)
by: Langford, Harry, et al.
Published: (2024)
LeakDojo: Decoding the Leakage Threats of RAG Systems
by: Zhang, Maosen, et al.
Published: (2026)
by: Zhang, Maosen, et al.
Published: (2026)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
Towards Understanding Unsafe Video Generation
by: Pang, Yan, et al.
Published: (2024)
by: Pang, Yan, et al.
Published: (2024)
Preserving Decision Sovereignty in Military AI: A Trade-Secret-Safe Architectural Framework for Model Replaceability, Human Authority, and State Control
by: Wei, Peng, et al.
Published: (2026)
by: Wei, Peng, et al.
Published: (2026)
Global Challenge for Safe and Secure LLMs Track 1
by: Jia, Xiaojun, et al.
Published: (2024)
by: Jia, Xiaojun, et al.
Published: (2024)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
by: Küchler, Nicolas, et al.
Published: (2025)
by: Küchler, Nicolas, et al.
Published: (2025)
CompLeak: Deep Learning Model Compression Exacerbates Privacy Leakage
by: Li, Na, et al.
Published: (2025)
by: Li, Na, et al.
Published: (2025)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
by: Dahiya, Pranav, et al.
Published: (2023)
by: Dahiya, Pranav, et al.
Published: (2023)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
Black-Box Access is Insufficient for Rigorous AI Audits
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
by: Lin, Zhiqiang, et al.
Published: (2025)
by: Lin, Zhiqiang, et al.
Published: (2025)
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025)
by: Barrett, Anthony M., et al.
Published: (2025)
Beyond Labeling Oracles: What does it mean to steal ML models?
by: Shafran, Avital, et al.
Published: (2023)
by: Shafran, Avital, et al.
Published: (2023)
Private, Verifiable, and Auditable AI Systems
by: South, Tobin
Published: (2025)
by: South, Tobin
Published: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
AI Propaganda factories with language models
by: Olejnik, Lukasz
Published: (2025)
by: Olejnik, Lukasz
Published: (2025)
Accelerating AI Development with Cyber Arenas
by: Cashman, William, et al.
Published: (2025)
by: Cashman, William, et al.
Published: (2025)
AI-Driven Cyber Threat Intelligence Automation
by: Shah, Shrit, et al.
Published: (2024)
by: Shah, Shrit, et al.
Published: (2024)
Securing the Future of GenAI: Policy and Technology
by: Christodorescu, Mihai, et al.
Published: (2024)
by: Christodorescu, Mihai, et al.
Published: (2024)
The Pitfalls of "Security by Obscurity" And What They Mean for Transparent AI
by: Hall, Peter, et al.
Published: (2025)
by: Hall, Peter, et al.
Published: (2025)
Coordinated Flaw Disclosure for AI: Beyond Security Vulnerabilities
by: Cattell, Sven, et al.
Published: (2024)
by: Cattell, Sven, et al.
Published: (2024)
Interplay of ISMS and AIMS in context of the EU AI Act
by: Pötsch, Jordan
Published: (2024)
by: Pötsch, Jordan
Published: (2024)
The End Of Universal Lifelong Identifiers: Identity Systems For The AI Era
by: Palakodety, Shriphani
Published: (2025)
by: Palakodety, Shriphani
Published: (2025)
Enabling External Scrutiny of AI Systems with Privacy-Enhancing Technologies
by: Beers, Kendrea, et al.
Published: (2025)
by: Beers, Kendrea, et al.
Published: (2025)
RedTeamLLM: an Agentic AI framework for offensive security
by: Challita, Brian, et al.
Published: (2025)
by: Challita, Brian, et al.
Published: (2025)
Governable AI: Provable Safety Under Extreme Threat Models
by: Wang, Donglin, et al.
Published: (2025)
by: Wang, Donglin, et al.
Published: (2025)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
Naming is framing: How cybersecurity's language problems are repeating in AI governance
by: Potter, Lianne
Published: (2025)
by: Potter, Lianne
Published: (2025)
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
Secure and Trustworthy Artificial Intelligence-Extended Reality (AI-XR) for Metaverses
by: Qayyum, Adnan, et al.
Published: (2022)
by: Qayyum, Adnan, et al.
Published: (2022)
Is Your AI Truly Yours? Leveraging Blockchain for Copyrights, Provenance, and Lineage
by: Wang, Qin, et al.
Published: (2024)
by: Wang, Qin, et al.
Published: (2024)
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
by: Feakins, Shaun, et al.
Published: (2026)
by: Feakins, Shaun, et al.
Published: (2026)
Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
by: Gressel, Gilad, et al.
Published: (2025)
by: Gressel, Gilad, et al.
Published: (2025)
Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation
by: Heiding, Fred, et al.
Published: (2025)
by: Heiding, Fred, et al.
Published: (2025)
Similar Items
-
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
by: Thudi, Anvith, et al.
Published: (2023) -
Architectural Neural Backdoors from First Principles
by: Langford, Harry, et al.
Published: (2024) -
LeakDojo: Decoding the Leakage Threats of RAG Systems
by: Zhang, Maosen, et al.
Published: (2026) -
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024) -
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)