Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMs
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Gupta, Aayush |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Age of Sensorial Zero Trust: Why We Can No Longer Trust Our Senses
par: Xavier, Fabio Correa
Publié: (2025)
par: Xavier, Fabio Correa
Publié: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
par: Dang, Kieu, et autres
Publié: (2025)
par: Dang, Kieu, et autres
Publié: (2025)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
par: Hill, Brennen, et autres
Publié: (2025)
par: Hill, Brennen, et autres
Publié: (2025)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
par: Nayak, Prabhudarshi, et autres
Publié: (2026)
par: Nayak, Prabhudarshi, et autres
Publié: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
par: Young, Richard J., et autres
Publié: (2026)
par: Young, Richard J., et autres
Publié: (2026)
Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation
par: Gharami, Kanchon, et autres
Publié: (2025)
par: Gharami, Kanchon, et autres
Publié: (2025)
BreakFun: Jailbreaking LLMs via Schema Exploitation
par: Oskooei, Amirkia Rafiei, et autres
Publié: (2025)
par: Oskooei, Amirkia Rafiei, et autres
Publié: (2025)
One-Shot Secure Aggregation: A Hybrid Cryptographic Protocol for Private Federated Learning in IoT
par: Emmaka, Imraul, et autres
Publié: (2025)
par: Emmaka, Imraul, et autres
Publié: (2025)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
par: Nathanson, Samuel, et autres
Publié: (2025)
par: Nathanson, Samuel, et autres
Publié: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
par: Othman, Refat
Publié: (2026)
par: Othman, Refat
Publié: (2026)
Non-Asymptotic Convergence of Discrete Diffusion Models: Masked and Random Walk dynamics
par: Conforti, Giovanni, et autres
Publié: (2025)
par: Conforti, Giovanni, et autres
Publié: (2025)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
par: Breneur, Oleksandr Marchenko, et autres
Publié: (2026)
par: Breneur, Oleksandr Marchenko, et autres
Publié: (2026)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
par: Tang, Zhengzheng
Publié: (2026)
par: Tang, Zhengzheng
Publié: (2026)
"Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
par: Piao, Yangheran, et autres
Publié: (2025)
par: Piao, Yangheran, et autres
Publié: (2025)
Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning
par: Huff, Philip, et autres
Publié: (2026)
par: Huff, Philip, et autres
Publié: (2026)
WAKESET: A Large-Scale, High-Reynolds Number Flow Dataset for Machine Learning of Turbulent Wake Dynamics
par: Cooper-Baldock, Zachary, et autres
Publié: (2026)
par: Cooper-Baldock, Zachary, et autres
Publié: (2026)
Scalpel-SAM: A Semi-Supervised Paradigm for Adapting SAM to Infrared Small Object Detection
par: Liu, Zihan, et autres
Publié: (2025)
par: Liu, Zihan, et autres
Publié: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
par: Asperti, Andrea, et autres
Publié: (2025)
par: Asperti, Andrea, et autres
Publié: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
par: Chakraborty, Amit, et autres
Publié: (2025)
par: Chakraborty, Amit, et autres
Publié: (2025)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
par: Zanbaghi, Shahin, et autres
Publié: (2025)
par: Zanbaghi, Shahin, et autres
Publié: (2025)
Seed-Induced Uniqueness in Transformer Models: Subspace Alignment Governs Subliminal Transfer
par: Okatan, Ayşe Selin, et autres
Publié: (2025)
par: Okatan, Ayşe Selin, et autres
Publié: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
par: Basu, Abhinaba
Publié: (2026)
par: Basu, Abhinaba
Publié: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
par: Giannini, Federico, et autres
Publié: (2026)
par: Giannini, Federico, et autres
Publié: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
par: Lasbordes, Maxence, et autres
Publié: (2026)
par: Lasbordes, Maxence, et autres
Publié: (2026)
The Automation Advantage in AI Red Teaming
par: Mulla, Rob, et autres
Publié: (2025)
par: Mulla, Rob, et autres
Publié: (2025)
Verifiable Fine-Tuning for LLMs: Zero-Knowledge Training Proofs Bound to Data Provenance and Policy
par: Akgul, Hasan, et autres
Publié: (2025)
par: Akgul, Hasan, et autres
Publié: (2025)
Unveiling Hidden Threats: Using Fractal Triggers to Boost Stealthiness of Distributed Backdoor Attacks in Federated Learning
par: Wang, Jian, et autres
Publié: (2025)
par: Wang, Jian, et autres
Publié: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
par: Dawson, Ads, et autres
Publié: (2025)
par: Dawson, Ads, et autres
Publié: (2025)
Pattern Recognition Tasks with Personalized Federated Learning
par: Rahman, Md. Arifur, et autres
Publié: (2026)
par: Rahman, Md. Arifur, et autres
Publié: (2026)
Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems
par: Parris, William
Publié: (2026)
par: Parris, William
Publié: (2026)
A Dual-Path Generative Framework for Zero-Day Fraud Detection in Banking Systems
par: Ismail, Nasim Abdirahman, et autres
Publié: (2026)
par: Ismail, Nasim Abdirahman, et autres
Publié: (2026)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
par: Collu, Matteo Gioele, et autres
Publié: (2023)
par: Collu, Matteo Gioele, et autres
Publié: (2023)
SALLIE: Safeguarding Against Latent Language & Image Exploits
par: Azov, Guy, et autres
Publié: (2026)
par: Azov, Guy, et autres
Publié: (2026)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
par: Ge, Yuxu
Publié: (2026)
par: Ge, Yuxu
Publié: (2026)
Towards Verifiable AI with Lightweight Cryptographic Proofs of Inference
par: Anchuri, Pranay, et autres
Publié: (2026)
par: Anchuri, Pranay, et autres
Publié: (2026)
Formal Analysis and Supply Chain Security for Agentic AI Skills
par: Bhardwaj, Varun Pratap
Publié: (2026)
par: Bhardwaj, Varun Pratap
Publié: (2026)
Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video
par: Moore, Alexander, et autres
Publié: (2025)
par: Moore, Alexander, et autres
Publié: (2025)
Harnessing non-adversarial robustness in large language models
par: Zhou, Qinghua, et autres
Publié: (2026)
par: Zhou, Qinghua, et autres
Publié: (2026)
Keys in the Weights: Transformer Authentication Using Model-Bound Latent Representations
par: Okatan, Ayşe S., et autres
Publié: (2025)
par: Okatan, Ayşe S., et autres
Publié: (2025)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
par: Gowda, Ishrith
Publié: (2026)
par: Gowda, Ishrith
Publié: (2026)
Documents similaires
-
The Age of Sensorial Zero Trust: Why We Can No Longer Trust Our Senses
par: Xavier, Fabio Correa
Publié: (2025) -
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
par: Dang, Kieu, et autres
Publié: (2025) -
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
par: Hill, Brennen, et autres
Publié: (2025) -
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
par: Nayak, Prabhudarshi, et autres
Publié: (2026) -
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
par: Young, Richard J., et autres
Publié: (2026)