Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
Fuente:
arXiv
Saved in:
| Main Author: | Sanna, Arun Chowdary |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
by: Lelle, Travis
Published: (2026)
by: Lelle, Travis
Published: (2026)
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
by: Hoang, Tien Dat
Published: (2025)
by: Hoang, Tien Dat
Published: (2025)
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
by: Pan, Zhixin, et al.
Published: (2025)
by: Pan, Zhixin, et al.
Published: (2025)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
by: Kumar, Sachin
Published: (2026)
by: Kumar, Sachin
Published: (2026)
SCAFDS: Edge-Feature Graph Attention for Interbank Fraud Detection with Attribution-Grounded SAR Generation
by: Uddin, Mohammad Nasir
Published: (2026)
by: Uddin, Mohammad Nasir
Published: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity
by: Rashidi, Mohammadreza
Published: (2026)
by: Rashidi, Mohammadreza
Published: (2026)
Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions
by: Ma, Jianan, et al.
Published: (2026)
by: Ma, Jianan, et al.
Published: (2026)
Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
by: Subedar, Noah, et al.
Published: (2025)
by: Subedar, Noah, et al.
Published: (2025)
Density-aware Sample-specific Attack
by: Wang, Qiyuan, et al.
Published: (2026)
by: Wang, Qiyuan, et al.
Published: (2026)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
by: Collu, Matteo Gioele, et al.
Published: (2023)
by: Collu, Matteo Gioele, et al.
Published: (2023)
Few-Shot Network Intrusion Detection Using Online Triplet Mining
by: Wilkie, Jack, et al.
Published: (2026)
by: Wilkie, Jack, et al.
Published: (2026)
A Novel Contrastive Loss for Zero-Day Network Intrusion Detection
by: Wilkie, Jack, et al.
Published: (2026)
by: Wilkie, Jack, et al.
Published: (2026)
A Protocol-Language Model for Network Intrusion (Without Deep Packet Inspection)
by: Sharma, Vivek Kumar
Published: (2026)
by: Sharma, Vivek Kumar
Published: (2026)
Contrastive Self-Supervised Network Intrusion Detection using Augmented Negative Pairs
by: Wilkie, Jack, et al.
Published: (2025)
by: Wilkie, Jack, et al.
Published: (2025)
A Dual-Path Generative Framework for Zero-Day Fraud Detection in Banking Systems
by: Ismail, Nasim Abdirahman, et al.
Published: (2026)
by: Ismail, Nasim Abdirahman, et al.
Published: (2026)
LLM Scalability Risk for Agentic-AI and Model Supply Chain Security
by: Ahi, Kiarash, et al.
Published: (2026)
by: Ahi, Kiarash, et al.
Published: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
A High-Recall Cost-Sensitive Machine Learning Framework for Real-Time Online Banking Transaction Fraud Detection
by: R., Karthikeyan V., et al.
Published: (2026)
by: R., Karthikeyan V., et al.
Published: (2026)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
Phishing Detection System: An Ensemble Approach Using Character-Level CNN and Feature Engineering
by: Dubey, Rudra, et al.
Published: (2025)
by: Dubey, Rudra, et al.
Published: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
by: Nathanson, Samuel, et al.
Published: (2025)
by: Nathanson, Samuel, et al.
Published: (2025)
From nuclear safety to LLM security: Applying non-probabilistic risk management strategies to build safe and secure LLM-powered systems
by: Gutfraind, Alexander, et al.
Published: (2025)
by: Gutfraind, Alexander, et al.
Published: (2025)
The Hiremath Early Detection (HED) Score: A Measure-Theoretic Evaluation Standard for Temporal Intelligence
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
Signal-Based Malware Classification Using 1D CNNs
by: Wilkie, Jack, et al.
Published: (2025)
by: Wilkie, Jack, et al.
Published: (2025)
Generalizable and Interpretable RF Fingerprinting with Shapelet-Enhanced Large Language Models
by: Zhao, Tianya, et al.
Published: (2026)
by: Zhao, Tianya, et al.
Published: (2026)
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
by: Pasupuleti, Vinil, et al.
Published: (2026)
by: Pasupuleti, Vinil, et al.
Published: (2026)
Is My Data in Your Retrieval Database? Membership Inference Attacks Against Retrieval Augmented Generation
by: Anderson, Maya, et al.
Published: (2024)
by: Anderson, Maya, et al.
Published: (2024)
Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions
by: Youssef, Michel
Published: (2025)
by: Youssef, Michel
Published: (2025)
Measuring Harmfulness of Computer-Using Agents
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
by: Tian, Aaron Xuxiang, et al.
Published: (2025)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
Send to which account? Evaluation of an LLM-based Scambaiting System
by: Siadati, Hossein, et al.
Published: (2025)
by: Siadati, Hossein, et al.
Published: (2025)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
by: Palaniappan, Nikitha M., et al.
Published: (2026)
by: Palaniappan, Nikitha M., et al.
Published: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
by: Xu, Luyao, et al.
Published: (2026)
by: Xu, Luyao, et al.
Published: (2026)
Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems
by: Cai, Yicheng, et al.
Published: (2026)
by: Cai, Yicheng, et al.
Published: (2026)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
by: Grofsky, Matthew
Published: (2025)
by: Grofsky, Matthew
Published: (2025)
Similar Items
-
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
by: Lelle, Travis
Published: (2026) -
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
by: Hoang, Tien Dat
Published: (2025) -
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
by: Pan, Zhixin, et al.
Published: (2025) -
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
by: Yu, Hao, et al.
Published: (2026) -
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
by: Kumar, Sachin
Published: (2026)