Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
Fuente:
arXiv
Guardado en:
| Autor principal: | Lelle, Travis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
por: Dang, Kieu, et al.
Publicado: (2025)
por: Dang, Kieu, et al.
Publicado: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
por: Azov, Guy, et al.
Publicado: (2026)
por: Azov, Guy, et al.
Publicado: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
por: Young, Richard J., et al.
Publicado: (2026)
por: Young, Richard J., et al.
Publicado: (2026)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
por: Gu, Yongtong, et al.
Publicado: (2026)
por: Gu, Yongtong, et al.
Publicado: (2026)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
por: Doda, Shravan
Publicado: (2026)
por: Doda, Shravan
Publicado: (2026)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
por: Othman, Refat
Publicado: (2026)
por: Othman, Refat
Publicado: (2026)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
por: Morales, Jaime, et al.
Publicado: (2026)
por: Morales, Jaime, et al.
Publicado: (2026)
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
por: Hoang, Tien Dat
Publicado: (2025)
por: Hoang, Tien Dat
Publicado: (2025)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
por: Wang, Haochuan Kevin, et al.
Publicado: (2026)
por: Wang, Haochuan Kevin, et al.
Publicado: (2026)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
por: Yu, Hao, et al.
Publicado: (2026)
por: Yu, Hao, et al.
Publicado: (2026)
Towards Modeling Cybersecurity Behavior of Humans in Organizations
por: Kürtz, Klaas Ole
Publicado: (2026)
por: Kürtz, Klaas Ole
Publicado: (2026)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
por: Hill, Brennen, et al.
Publicado: (2025)
por: Hill, Brennen, et al.
Publicado: (2025)
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
por: Sanna, Arun Chowdary
Publicado: (2025)
por: Sanna, Arun Chowdary
Publicado: (2025)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
por: Dawson, Ads, et al.
Publicado: (2025)
por: Dawson, Ads, et al.
Publicado: (2025)
The Automation Advantage in AI Red Teaming
por: Mulla, Rob, et al.
Publicado: (2025)
por: Mulla, Rob, et al.
Publicado: (2025)
Retrieval Augmented Classification for Confidential Documents
por: Chang, Yeseul E., et al.
Publicado: (2026)
por: Chang, Yeseul E., et al.
Publicado: (2026)
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
por: Young, Richard J., et al.
Publicado: (2026)
por: Young, Richard J., et al.
Publicado: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
por: Hiremath, Prakul Sunil, et al.
Publicado: (2026)
por: Hiremath, Prakul Sunil, et al.
Publicado: (2026)
Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
por: Adjei, Yaw Osei, et al.
Publicado: (2025)
por: Adjei, Yaw Osei, et al.
Publicado: (2025)
Send to which account? Evaluation of an LLM-based Scambaiting System
por: Siadati, Hossein, et al.
Publicado: (2025)
por: Siadati, Hossein, et al.
Publicado: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
por: Chona, Alankrit, et al.
Publicado: (2026)
por: Chona, Alankrit, et al.
Publicado: (2026)
Measuring Harmfulness of Computer-Using Agents
por: Tian, Aaron Xuxiang, et al.
Publicado: (2025)
por: Tian, Aaron Xuxiang, et al.
Publicado: (2025)
Countermind: A Multi-Layered Security Architecture for Large Language Models
por: Schwarz, Dominik
Publicado: (2025)
por: Schwarz, Dominik
Publicado: (2025)
Density-aware Sample-specific Attack
por: Wang, Qiyuan, et al.
Publicado: (2026)
por: Wang, Qiyuan, et al.
Publicado: (2026)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
por: Nathanson, Samuel, et al.
Publicado: (2025)
por: Nathanson, Samuel, et al.
Publicado: (2025)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
por: Zanbaghi, Shahin, et al.
Publicado: (2025)
por: Zanbaghi, Shahin, et al.
Publicado: (2025)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
por: Seddik, Issam, et al.
Publicado: (2025)
por: Seddik, Issam, et al.
Publicado: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
por: Chakraborty, Amit, et al.
Publicado: (2025)
por: Chakraborty, Amit, et al.
Publicado: (2025)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
por: Palaniappan, Nikitha M., et al.
Publicado: (2026)
por: Palaniappan, Nikitha M., et al.
Publicado: (2026)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
por: Grofsky, Matthew
Publicado: (2025)
por: Grofsky, Matthew
Publicado: (2025)
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
por: Pan, Zhixin, et al.
Publicado: (2025)
por: Pan, Zhixin, et al.
Publicado: (2025)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
por: Ge, Yuxu
Publicado: (2026)
por: Ge, Yuxu
Publicado: (2026)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
por: Khatiwala, Jeel Piyushkumar, et al.
Publicado: (2026)
por: Khatiwala, Jeel Piyushkumar, et al.
Publicado: (2026)
Towards Agentic Investigation of Security Alerts
por: Eilertsen, Even, et al.
Publicado: (2026)
por: Eilertsen, Even, et al.
Publicado: (2026)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
por: Zhang, Tian, et al.
Publicado: (2026)
por: Zhang, Tian, et al.
Publicado: (2026)
Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
por: Subedar, Noah, et al.
Publicado: (2025)
por: Subedar, Noah, et al.
Publicado: (2025)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
por: Leonesi, Matteo, et al.
Publicado: (2026)
por: Leonesi, Matteo, et al.
Publicado: (2026)
Phishing Detection System: An Ensemble Approach Using Character-Level CNN and Feature Engineering
por: Dubey, Rudra, et al.
Publicado: (2025)
por: Dubey, Rudra, et al.
Publicado: (2025)
SCAFDS: Edge-Feature Graph Attention for Interbank Fraud Detection with Attribution-Grounded SAR Generation
por: Uddin, Mohammad Nasir
Publicado: (2026)
por: Uddin, Mohammad Nasir
Publicado: (2026)
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
por: Usman, Rana Muhammad
Publicado: (2026)
por: Usman, Rana Muhammad
Publicado: (2026)
Ejemplares similares
-
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
por: Dang, Kieu, et al.
Publicado: (2025) -
SALLIE: Safeguarding Against Latent Language & Image Exploits
por: Azov, Guy, et al.
Publicado: (2026) -
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
por: Young, Richard J., et al.
Publicado: (2026) -
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
por: Gu, Yongtong, et al.
Publicado: (2026) -
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
por: Doda, Shravan
Publicado: (2026)