Attacking interpretable NLP systems
Fuente:
arXiv
Saved in:
| Main Authors: | Abdukhamidov, Eldor, Abuhmed, Tamer, Santos, Joanna C. S., Abuhamad, Mohammed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
by: Abdukhamidov, Eldor, et al.
Published: (2025)
by: Abdukhamidov, Eldor, et al.
Published: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026)
by: Othman, Refat
Published: (2026)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
by: Gowda, Ishrith
Published: (2026)
by: Gowda, Ishrith
Published: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
by: Bommarito II, Michael J.
Published: (2025)
by: Bommarito II, Michael J.
Published: (2025)
Lightweight LLMs for Network Attack Detection in IoT Networks
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026)
by: Sudasinghe, Piyumi Bhagya, et al.
Published: (2026)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
by: Nathanson, Samuel, et al.
Published: (2025)
by: Nathanson, Samuel, et al.
Published: (2025)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
by: Rosenblatt, Lucas, et al.
Published: (2026)
by: Rosenblatt, Lucas, et al.
Published: (2026)
Detecting Prompt Injection Attacks Against Application Using Classifiers
by: Shaheer, Safwan, et al.
Published: (2025)
by: Shaheer, Safwan, et al.
Published: (2025)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
by: Shaheer, Safwan, et al.
Published: (2025)
by: Shaheer, Safwan, et al.
Published: (2025)
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
by: Pulipaka, Sidharth, et al.
Published: (2026)
by: Pulipaka, Sidharth, et al.
Published: (2026)
Training Language Models to Use Prolog as a Tool
by: Mellgren, Niklas, et al.
Published: (2025)
by: Mellgren, Niklas, et al.
Published: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
by: Dawson, Ads, et al.
Published: (2025)
by: Dawson, Ads, et al.
Published: (2025)
The Automation Advantage in AI Red Teaming
by: Mulla, Rob, et al.
Published: (2025)
by: Mulla, Rob, et al.
Published: (2025)
Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions
by: Youssef, Michel
Published: (2025)
by: Youssef, Michel
Published: (2025)
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
by: Cotti, Luca, et al.
Published: (2025)
by: Cotti, Luca, et al.
Published: (2025)
Unsupervised Baseline Clustering and Incremental Adaptation for IoT Device Traffic Profiling
by: Alderman, Sean M., et al.
Published: (2026)
by: Alderman, Sean M., et al.
Published: (2026)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
by: Ge, Yuxu
Published: (2026)
by: Ge, Yuxu
Published: (2026)
Poison in the Well: Feature Embedding Disruption in Backdoor Attacks
by: Feng, Zhou, et al.
Published: (2025)
by: Feng, Zhou, et al.
Published: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
by: Seddik, Issam, et al.
Published: (2025)
by: Seddik, Issam, et al.
Published: (2025)
Enhancing Large Language Models through Neuro-Symbolic Integration and Ontological Reasoning
by: Vsevolodovna, Ruslan Idelfonso Magana, et al.
Published: (2025)
by: Vsevolodovna, Ruslan Idelfonso Magana, et al.
Published: (2025)
Security Considerations for Multi-agent Systems
by: Nguyen, Tam, et al.
Published: (2026)
by: Nguyen, Tam, et al.
Published: (2026)
Software Defined Vehicle Code Generation: A Few-Shot Prompting Approach
by: Nguyen, Quang-Dung, et al.
Published: (2025)
by: Nguyen, Quang-Dung, et al.
Published: (2025)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
by: Bercovich, Ivan, et al.
Published: (2026)
by: Bercovich, Ivan, et al.
Published: (2026)
Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property Refinement
by: Ma, Jianan, et al.
Published: (2025)
by: Ma, Jianan, et al.
Published: (2025)
Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation
by: Gharami, Kanchon, et al.
Published: (2025)
by: Gharami, Kanchon, et al.
Published: (2025)
LLMCup: Ranking-Enhanced Comment Updating with LLMs
by: Ge, Hua, et al.
Published: (2025)
by: Ge, Hua, et al.
Published: (2025)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
by: Grofsky, Matthew
Published: (2025)
by: Grofsky, Matthew
Published: (2025)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
by: Khatiwala, Jeel Piyushkumar, et al.
Published: (2026)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
by: Palaniappan, Nikitha M., et al.
Published: (2026)
by: Palaniappan, Nikitha M., et al.
Published: (2026)
Towards Agentic Investigation of Security Alerts
by: Eilertsen, Even, et al.
Published: (2026)
by: Eilertsen, Even, et al.
Published: (2026)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026)
by: Leonesi, Matteo, et al.
Published: (2026)
Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMs
by: Gupta, Aayush
Published: (2025)
by: Gupta, Aayush
Published: (2025)
ZK-SenseLM: Verifiable Large-Model Wireless Sensing with Selective Abstention and Zero-Knowledge Attestation
by: Akgul, Hasan, et al.
Published: (2025)
by: Akgul, Hasan, et al.
Published: (2025)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025)
by: Williams-King, David, et al.
Published: (2025)
Quantifying Return on Security Controls in LLM Systems
by: Moulton, Richard Helder, et al.
Published: (2025)
by: Moulton, Richard Helder, et al.
Published: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs
by: Bakker, Isabelle, et al.
Published: (2025)
by: Bakker, Isabelle, et al.
Published: (2025)
Similar Items
-
Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
by: Abdukhamidov, Eldor, et al.
Published: (2025) -
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026) -
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
by: Gowda, Ishrith
Published: (2026) -
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026) -
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
by: Bommarito II, Michael J.
Published: (2025)