Narrow Secret Loyalty Dodges Black-Box Audits
Fuente:
arXiv
Saved in:
| Main Authors: | Lamerton, Alfie, Roger, Fabien |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Black-Box Access is Insufficient for Rigorous AI Audits
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
Comparative Analysis of Black-Box and White-Box Machine Learning Model in Phishing Detection
by: Fajar, Abdullah, et al.
Published: (2024)
by: Fajar, Abdullah, et al.
Published: (2024)
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
by: Domico, Kyle, et al.
Published: (2025)
by: Domico, Kyle, et al.
Published: (2025)
FDLLM: A Dedicated Detector for Black-Box LLMs Fingerprinting
by: Fu, Zhiyuan, et al.
Published: (2025)
by: Fu, Zhiyuan, et al.
Published: (2025)
Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software
by: Kordonsky, Tomer, et al.
Published: (2026)
by: Kordonsky, Tomer, et al.
Published: (2026)
Vulnerability Disclosure through Adaptive Black-Box Adversarial Attacks on NIDS
by: Ennaji, Sabrine, et al.
Published: (2025)
by: Ennaji, Sabrine, et al.
Published: (2025)
Turning Black Box into White Box: Dataset Distillation Leaks
by: Chen, Huajie, et al.
Published: (2026)
by: Chen, Huajie, et al.
Published: (2026)
Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
by: Li, Yakai, et al.
Published: (2025)
by: Li, Yakai, et al.
Published: (2025)
How stealthy is stealthy? Studying the Efficacy of Black-Box Adversarial Attacks in the Real World
by: Panebianco, Francesco, et al.
Published: (2025)
by: Panebianco, Francesco, et al.
Published: (2025)
E-MIA: Exam-Style Black-Box Membership Inference Attacks against RAG Systems
by: Guan, Zelin, et al.
Published: (2026)
by: Guan, Zelin, et al.
Published: (2026)
Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS
by: Ennaji, Sabrine, et al.
Published: (2025)
by: Ennaji, Sabrine, et al.
Published: (2025)
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
by: Cheng, Ruoxi, et al.
Published: (2024)
by: Cheng, Ruoxi, et al.
Published: (2024)
Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models
by: Yang, Yiqi, et al.
Published: (2024)
by: Yang, Yiqi, et al.
Published: (2024)
Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing
by: Yin, Jinhua, et al.
Published: (2025)
by: Yin, Jinhua, et al.
Published: (2025)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
by: Wang, Haowei, et al.
Published: (2025)
by: Wang, Haowei, et al.
Published: (2025)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
by: Najt, Elle, et al.
Published: (2026)
by: Najt, Elle, et al.
Published: (2026)
ArmSSL: Adversarial Robust Black-Box Watermarking for Self-Supervised Learning Pre-trained Encoders
by: Jiang, Yongqi, et al.
Published: (2026)
by: Jiang, Yongqi, et al.
Published: (2026)
RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models
by: Lv, Peizhuo, et al.
Published: (2025)
by: Lv, Peizhuo, et al.
Published: (2025)
StruPhantom: Evolutionary Injection Attacks on Black-Box Tabular Agents Powered by Large Language Models
by: Feng, Yang, et al.
Published: (2025)
by: Feng, Yang, et al.
Published: (2025)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-supervised Learning
by: Lv, Peizhuo, et al.
Published: (2022)
by: Lv, Peizhuo, et al.
Published: (2022)
Symmetry Defeats Auditing
by: Merrill, Nick, et al.
Published: (2026)
by: Merrill, Nick, et al.
Published: (2026)
Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models
by: Wang, Shuqiang, et al.
Published: (2026)
by: Wang, Shuqiang, et al.
Published: (2026)
Tricking LLM-Based NPCs into Spilling Secrets
by: Shiomi, Kyohei, et al.
Published: (2025)
by: Shiomi, Kyohei, et al.
Published: (2025)
SUDP: Secret-Use Delegation Protocol for Agentic Systems
by: Yu, Xiaohang, et al.
Published: (2026)
by: Yu, Xiaohang, et al.
Published: (2026)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026)
by: Lu, Yu-An, et al.
Published: (2026)
FTSmartAudit: A Knowledge Distillation-Enhanced Framework for Automated Smart Contract Auditing Using Fine-Tuned LLMs
by: Wei, Zhiyuan, et al.
Published: (2024)
by: Wei, Zhiyuan, et al.
Published: (2024)
On the Evidentiary Limits of Membership Inference for Copyright Auditing
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
Detecting Adversarial Fine-tuning with Auditing Agents
by: Egler, Sarah, et al.
Published: (2025)
by: Egler, Sarah, et al.
Published: (2025)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
by: Chen, Meifang, et al.
Published: (2026)
by: Chen, Meifang, et al.
Published: (2026)
CapSeal: Capability-Sealed Secret Mediation for Secure Agent Execution
by: Jin, Shutong, et al.
Published: (2026)
by: Jin, Shutong, et al.
Published: (2026)
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
by: Li, Xinyu, et al.
Published: (2025)
by: Li, Xinyu, et al.
Published: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
by: Gohil, Vasudev
Published: (2025)
by: Gohil, Vasudev
Published: (2025)
When Agents Handle Secrets: A Survey of Confidential Computing for Agentic AI
by: Forough, Javad, et al.
Published: (2026)
by: Forough, Javad, et al.
Published: (2026)
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
An Empirical Audit of k-NAF Budget Accounting for Anchored Decoding
by: Vijayavallabh, J.
Published: (2026)
by: Vijayavallabh, J.
Published: (2026)
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
by: Lv, Lijia, et al.
Published: (2026)
by: Lv, Lijia, et al.
Published: (2026)
Similar Items
-
Black-Box Access is Insufficient for Rigorous AI Audits
by: Casper, Stephen, et al.
Published: (2024) -
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
by: Zhu, Xiaoyuan, et al.
Published: (2025) -
Comparative Analysis of Black-Box and White-Box Machine Learning Model in Phishing Detection
by: Fajar, Abdullah, et al.
Published: (2024) -
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
by: Domico, Kyle, et al.
Published: (2025) -
FDLLM: A Dedicated Detector for Black-Box LLMs Fingerprinting
by: Fu, Zhiyuan, et al.
Published: (2025)