Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection
Fuente:
arXiv
Saved in:
| Main Author: | Graves, Marcus |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
Send to which account? Evaluation of an LLM-based Scambaiting System
by: Siadati, Hossein, et al.
Published: (2025)
by: Siadati, Hossein, et al.
Published: (2025)
From nuclear safety to LLM security: Applying non-probabilistic risk management strategies to build safe and secure LLM-powered systems
by: Gutfraind, Alexander, et al.
Published: (2025)
by: Gutfraind, Alexander, et al.
Published: (2025)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
Active Authentication via Korean Keystrokes Under Varying LLM Assistance and Cognitive Contexts
by: Roh, Dong Hyun, et al.
Published: (2025)
by: Roh, Dong Hyun, et al.
Published: (2025)
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
by: Pan, Zhixin, et al.
Published: (2025)
by: Pan, Zhixin, et al.
Published: (2025)
Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection
by: Corll, J Alex
Published: (2026)
by: Corll, J Alex
Published: (2026)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems
by: Cai, Yicheng, et al.
Published: (2026)
by: Cai, Yicheng, et al.
Published: (2026)
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
by: Sanna, Arun Chowdary
Published: (2025)
by: Sanna, Arun Chowdary
Published: (2025)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
Predicting known Vulnerabilities from Attack News: A Transformer-Based Approach
by: Othman, Refat, et al.
Published: (2026)
by: Othman, Refat, et al.
Published: (2026)
The Vehicle May Be Sick: Denial of Diagnostic Services by Exploiting the CAN Transport Protocol
by: Baek, Seungjin, et al.
Published: (2026)
by: Baek, Seungjin, et al.
Published: (2026)
Enhancing Energy Sector Resilience: Integrating Security by Design Principles
by: Shirtz, Dov, et al.
Published: (2024)
by: Shirtz, Dov, et al.
Published: (2024)
Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced Attack
by: Chen, Guanzhong, et al.
Published: (2024)
by: Chen, Guanzhong, et al.
Published: (2024)
ADMIn: Attacks on Dataset, Model and Input. A Threat Model for AI Based Software
by: Kumar, Vimal, et al.
Published: (2024)
by: Kumar, Vimal, et al.
Published: (2024)
Optimisation of cyber insurance coverage with selection of cost effective security controls
by: Uuganbayar, Ganbayar, et al.
Published: (2025)
by: Uuganbayar, Ganbayar, et al.
Published: (2025)
A Relevance Model for Threat-Centric Ranking of Cybersecurity Vulnerabilities
by: McCoy, Corren, et al.
Published: (2024)
by: McCoy, Corren, et al.
Published: (2024)
Rubber Mallet: A Study of High Frequency Localized Bit Flips and Their Impact on Security
by: Adiletta, Andrew, et al.
Published: (2025)
by: Adiletta, Andrew, et al.
Published: (2025)
I Know What You Said: Unveiling Hardware Cache Side-Channels in Local Large Language Model Inference
by: Gao, Zibo, et al.
Published: (2025)
by: Gao, Zibo, et al.
Published: (2025)
Criminal Liability in AI-Enabled Autonomous Vehicles: A Comparative Study
by: Singh, Sahibpreet, et al.
Published: (2025)
by: Singh, Sahibpreet, et al.
Published: (2025)
Cybercrime and Computer Forensics in Epoch of Artificial Intelligence in India
by: Singh, Sahibpreet, et al.
Published: (2025)
by: Singh, Sahibpreet, et al.
Published: (2025)
Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
by: Xu, Luyao, et al.
Published: (2026)
by: Xu, Luyao, et al.
Published: (2026)
Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats
by: Zhang, Bingxue, et al.
Published: (2026)
by: Zhang, Bingxue, et al.
Published: (2026)
Measuring Harmfulness of Computer-Using Agents
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
Countermind: A Multi-Layered Security Architecture for Large Language Models
by: Schwarz, Dominik
Published: (2025)
by: Schwarz, Dominik
Published: (2025)
Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming
by: Janjusevic, Strahinja, et al.
Published: (2025)
by: Janjusevic, Strahinja, et al.
Published: (2025)
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity
by: Rashidi, Mohammadreza
Published: (2026)
by: Rashidi, Mohammadreza
Published: (2026)
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
by: Muzsai, Lajos, et al.
Published: (2024)
by: Muzsai, Lajos, et al.
Published: (2024)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
False Security Confidence in Benign LLM Code Generation
by: Ren, Xiaolei
Published: (2026)
by: Ren, Xiaolei
Published: (2026)
Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions
by: Ma, Jianan, et al.
Published: (2026)
by: Ma, Jianan, et al.
Published: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
by: Kim, Minseok, et al.
Published: (2025)
by: Kim, Minseok, et al.
Published: (2025)
RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code
by: Pellew, John, et al.
Published: (2026)
by: Pellew, John, et al.
Published: (2026)
Is My Data in Your Retrieval Database? Membership Inference Attacks Against Retrieval Augmented Generation
by: Anderson, Maya, et al.
Published: (2024)
by: Anderson, Maya, et al.
Published: (2024)
CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
by: Muzsai, Lajos, et al.
Published: (2025)
by: Muzsai, Lajos, et al.
Published: (2025)
ILION: Deterministic Pre-Execution Safety Gates for Agentic AI Systems
by: Chitan, Florin Adrian
Published: (2026)
by: Chitan, Florin Adrian
Published: (2026)
AITH: A Post-Quantum Continuous Delegation Protocol for Human-AI Trust Establishment
by: Chen, Zhaoliang
Published: (2026)
by: Chen, Zhaoliang
Published: (2026)
Similar Items
-
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
by: Adiletta, Andrew, et al.
Published: (2025) -
Send to which account? Evaluation of an LLM-based Scambaiting System
by: Siadati, Hossein, et al.
Published: (2025) -
From nuclear safety to LLM security: Applying non-probabilistic risk management strategies to build safe and secure LLM-powered systems
by: Gutfraind, Alexander, et al.
Published: (2025) -
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026) -
Active Authentication via Korean Keystrokes Under Varying LLM Assistance and Cognitive Contexts
by: Roh, Dong Hyun, et al.
Published: (2025)