Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Davies, Xander, Winsor, Eric, Souly, Alexandra, Korbak, Tomek, Kirk, Robert, de Witt, Christian Schroeder, Gal, Yarin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UK AISI Alignment Evaluation Case-Study
von: Souly, Alexandra, et al.
Veröffentlicht: (2026)
von: Souly, Alexandra, et al.
Veröffentlicht: (2026)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
A sketch of an AI control safety case
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)
Extending the OWASP Multi-Agentic System Threat Modeling Guide: Insights from Multi-Agent Security Research
von: Krawiecka, Klaudia, et al.
Veröffentlicht: (2025)
von: Krawiecka, Klaudia, et al.
Veröffentlicht: (2025)
Safety case template for frontier AI: A cyber inability argument
von: Goemans, Arthur, et al.
Veröffentlicht: (2024)
von: Goemans, Arthur, et al.
Veröffentlicht: (2024)
Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations
von: Del Rosario, Ron F., et al.
Veröffentlicht: (2025)
von: Del Rosario, Ron F., et al.
Veröffentlicht: (2025)
Log Probability Tracking of LLM APIs
von: Chauvin, Timothée, et al.
Veröffentlicht: (2025)
von: Chauvin, Timothée, et al.
Veröffentlicht: (2025)
Adaptive Exploit Generation against Security Devices and Security APIs
von: Künnemann, Robert, et al.
Veröffentlicht: (2024)
von: Künnemann, Robert, et al.
Veröffentlicht: (2024)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
von: Sneh, Jonathan, et al.
Veröffentlicht: (2025)
von: Sneh, Jonathan, et al.
Veröffentlicht: (2025)
Practical challenges of control monitoring in frontier AI deployments
von: Lindner, David, et al.
Veröffentlicht: (2025)
von: Lindner, David, et al.
Veröffentlicht: (2025)
Token-Efficient Change Detection in LLM APIs
von: Chauvin, Timothée, et al.
Veröffentlicht: (2026)
von: Chauvin, Timothée, et al.
Veröffentlicht: (2026)
VET Your Agent: Towards Host-Independent Autonomy via Verifiable Execution Traces
von: Grigor, Artem, et al.
Veröffentlicht: (2025)
von: Grigor, Artem, et al.
Veröffentlicht: (2025)
API Security Based on Automatic OpenAPI Mapping
von: Levi, Yarin, et al.
Veröffentlicht: (2026)
von: Levi, Yarin, et al.
Veröffentlicht: (2026)
AEX: Non-Intrusive Multi-Hop Attestation and Provenance for LLM APIs
von: Guan, Yongjie
Veröffentlicht: (2026)
von: Guan, Yongjie
Veröffentlicht: (2026)
Mining REST APIs for Potential Mass Assignment Vulnerabilities
von: Mazidi, Arash, et al.
Veröffentlicht: (2024)
von: Mazidi, Arash, et al.
Veröffentlicht: (2024)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
von: Che, Zora, et al.
Veröffentlicht: (2025)
von: Che, Zora, et al.
Veröffentlicht: (2025)
On Digital Twins in Defence: Overview and Applications
von: Giberna, Marco, et al.
Veröffentlicht: (2025)
von: Giberna, Marco, et al.
Veröffentlicht: (2025)
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
von: Zeng, Xinyi, et al.
Veröffentlicht: (2024)
von: Zeng, Xinyi, et al.
Veröffentlicht: (2024)
MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
von: Aichberger, Lukas, et al.
Veröffentlicht: (2025)
von: Aichberger, Lukas, et al.
Veröffentlicht: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
MirrorFuzz: Leveraging LLM and Shared Bugs for Deep Learning Framework APIs Fuzzing
von: Ou, Shiwen, et al.
Veröffentlicht: (2025)
von: Ou, Shiwen, et al.
Veröffentlicht: (2025)
Architecture Matters for Multi-Agent Security
von: Hagag, Ben, et al.
Veröffentlicht: (2026)
von: Hagag, Ben, et al.
Veröffentlicht: (2026)
A Complexity-Informed Approach to Optimise Cyber Defences
von: Alevizos, Lampis
Veröffentlicht: (2025)
von: Alevizos, Lampis
Veröffentlicht: (2025)
R+R: Reassessing Java Security API Misuse in Current LLMs: A Replication on JCA and JSSE APIs with External Security Knowledge
von: Lu, Tianhe, et al.
Veröffentlicht: (2026)
von: Lu, Tianhe, et al.
Veröffentlicht: (2026)
Computing Low-Entropy Couplings for Large-Support Distributions
von: Sokota, Samuel, et al.
Veröffentlicht: (2024)
von: Sokota, Samuel, et al.
Veröffentlicht: (2024)
Multi-Objective Reinforcement Learning for Automated Resilient Cyber Defence
von: O'Driscoll, Ross, et al.
Veröffentlicht: (2024)
von: O'Driscoll, Ross, et al.
Veröffentlicht: (2024)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
All Your Knowledge Belongs to Us: Stealing Knowledge Graphs via Reasoning APIs
von: Xi, Zhaohan
Veröffentlicht: (2025)
von: Xi, Zhaohan
Veröffentlicht: (2025)
Unelicitable Backdoors in Language Models via Cryptographic Transformer Circuits
von: Draguns, Andis, et al.
Veröffentlicht: (2024)
von: Draguns, Andis, et al.
Veröffentlicht: (2024)
PSyDUCK: Training-Free Steganography for Latent Diffusion
von: Mahfuz, Aqib, et al.
Veröffentlicht: (2025)
von: Mahfuz, Aqib, et al.
Veröffentlicht: (2025)
The Fundamental Limits of Least-Privilege Learning
von: Stadler, Theresa, et al.
Veröffentlicht: (2024)
von: Stadler, Theresa, et al.
Veröffentlicht: (2024)
Paladin: A Policy Framework for Securing Cloud APIs by Combining Application Context with Generative AI
von: Priya, Shriti, et al.
Veröffentlicht: (2026)
von: Priya, Shriti, et al.
Veröffentlicht: (2026)
Deep Dive into the Abuse of DL APIs To Create Malicious AI Models and How to Detect Them
von: Nabeel, Mohamed, et al.
Veröffentlicht: (2026)
von: Nabeel, Mohamed, et al.
Veröffentlicht: (2026)
Analysing India's Cyber Warfare Readiness and Developing a Defence Strategy
von: Fernandes, Yohan, et al.
Veröffentlicht: (2024)
von: Fernandes, Yohan, et al.
Veröffentlicht: (2024)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
von: Foerster, Hanna, et al.
Veröffentlicht: (2025)
von: Foerster, Hanna, et al.
Veröffentlicht: (2025)
Semantic Denial of Service in LLM-controlled robots
von: Steinberg, Jonathan, et al.
Veröffentlicht: (2026)
von: Steinberg, Jonathan, et al.
Veröffentlicht: (2026)
Detecting Misuse of Security APIs: A Systematic Review
von: Mousavi, Zahra, et al.
Veröffentlicht: (2023)
von: Mousavi, Zahra, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
UK AISI Alignment Evaluation Case-Study
von: Souly, Alexandra, et al.
Veröffentlicht: (2026) -
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
von: Korbak, Tomek, et al.
Veröffentlicht: (2025) -
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
von: Bazinska, Julia, et al.
Veröffentlicht: (2025) -
A sketch of an AI control safety case
von: Korbak, Tomek, et al.
Veröffentlicht: (2025) -
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)