Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
Fuente:
arXiv
Saved in:
| Main Authors: | Chia, Xin Wei, Wong, Swee Liang, Pan, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt Inject Detection with Generative Explanation as an Investigative Tool
by: Pan, Jonathan, et al.
Published: (2025)
by: Pan, Jonathan, et al.
Published: (2025)
Enhancing Reasoning Capacity of SLM using Cognitive Enhancement
by: Pan, Jonathan, et al.
Published: (2024)
by: Pan, Jonathan, et al.
Published: (2024)
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
by: Singh, Inderjeet, et al.
Published: (2026)
by: Singh, Inderjeet, et al.
Published: (2026)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
by: Liang, Buyun, et al.
Published: (2026)
by: Liang, Buyun, et al.
Published: (2026)
Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
ADR: An Agentic Detection System for Enterprise Agentic AI Security
by: Li, Chenning, et al.
Published: (2026)
by: Li, Chenning, et al.
Published: (2026)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Are Robust LLM Fingerprints Adversarially Robust?
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
by: Yang, Zhuoran, et al.
Published: (2026)
by: Yang, Zhuoran, et al.
Published: (2026)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses
by: Yichao, Wu, et al.
Published: (2025)
by: Yichao, Wu, et al.
Published: (2025)
Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
Enhancing Reliability in LLM-Based Secure Code Generation
by: Kharma, Mohammed F., et al.
Published: (2026)
by: Kharma, Mohammed F., et al.
Published: (2026)
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
by: Sun, Luze, et al.
Published: (2026)
by: Sun, Luze, et al.
Published: (2026)
EVMbench: Evaluating AI Agents on Smart Contract Security
by: Wang, Justin, et al.
Published: (2026)
by: Wang, Justin, et al.
Published: (2026)
SoK: Security and Privacy Risks of Healthcare AI
by: Chang, Yuanhaur, et al.
Published: (2024)
by: Chang, Yuanhaur, et al.
Published: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
by: Yang, Xianglin, et al.
Published: (2026)
by: Yang, Xianglin, et al.
Published: (2026)
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
Runtime Detection of Adversarial Attacks in AI Accelerators Using Performance Counters
by: Rahaman, Habibur, et al.
Published: (2025)
by: Rahaman, Habibur, et al.
Published: (2025)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
by: Panfilov, Alexander, et al.
Published: (2026)
by: Panfilov, Alexander, et al.
Published: (2026)
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
by: Kharma, Mohammed, et al.
Published: (2026)
by: Kharma, Mohammed, et al.
Published: (2026)
Smart Water Security with AI and Blockchain-Enhanced Digital Twins
by: Homaei, Mohammadhossein, et al.
Published: (2025)
by: Homaei, Mohammadhossein, et al.
Published: (2025)
SAGA: A Security Architecture for Governing AI Agentic Systems
by: Syros, Georgios, et al.
Published: (2025)
by: Syros, Georgios, et al.
Published: (2025)
Position: AI Security Policy Should Target Systems, Not Models
by: Riegler, Michael A., et al.
Published: (2026)
by: Riegler, Michael A., et al.
Published: (2026)
Trust No AI: Prompt Injection Along The CIA Security Triad
by: Rehberger, Johann
Published: (2024)
by: Rehberger, Johann
Published: (2024)
Trustworthy Blockchain-based Federated Learning for Electronic Health Records: Securing Participant Identity with Decentralized Identifiers and Verifiable Credentials
by: Tertulino, Rodrigo, et al.
Published: (2026)
by: Tertulino, Rodrigo, et al.
Published: (2026)
The Dark Side of Digital Twins: Adversarial Attacks on AI-Driven Water Forecasting
by: Homaei, Mohammadhossein, et al.
Published: (2025)
by: Homaei, Mohammadhossein, et al.
Published: (2025)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
by: Bazinska, Julia, et al.
Published: (2025)
by: Bazinska, Julia, et al.
Published: (2025)
CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
by: Das, Debeshee, et al.
Published: (2025)
by: Das, Debeshee, et al.
Published: (2025)
Privacy and Security Implications of Cloud-Based AI Services : A Survey
by: Luqman, Alka, et al.
Published: (2024)
by: Luqman, Alka, et al.
Published: (2024)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
by: Wu, Yuanwei, et al.
Published: (2023)
by: Wu, Yuanwei, et al.
Published: (2023)
Taming Data Challenges in ML-based Security Tasks Using Generative AI
by: Kanchi, Shravya, et al.
Published: (2025)
by: Kanchi, Shravya, et al.
Published: (2025)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
Black-box Adversarial Attacks on Network-wide Multi-step Traffic State Prediction Models
by: Poudel, Bibek, et al.
Published: (2021)
by: Poudel, Bibek, et al.
Published: (2021)
A General Black-box Adversarial Attack on Graph-based Fake News Detectors
by: Zhu, Peican, et al.
Published: (2024)
by: Zhu, Peican, et al.
Published: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024)
by: Łucki, Jakub, et al.
Published: (2024)
Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
by: Maiorano, Alexandre Cristovão
Published: (2026)
by: Maiorano, Alexandre Cristovão
Published: (2026)
Similar Items
-
Prompt Inject Detection with Generative Explanation as an Investigative Tool
by: Pan, Jonathan, et al.
Published: (2025) -
Enhancing Reasoning Capacity of SLM using Cognitive Enhancement
by: Pan, Jonathan, et al.
Published: (2024) -
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
by: Singh, Inderjeet, et al.
Published: (2026) -
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
by: Liang, Buyun, et al.
Published: (2026) -
Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
by: Thornton, Scott
Published: (2026)