One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
Fuente:
arXiv
Guardado en:
| Autores principales: | Zloczower, Itay, Lenga, Eyal, Gressel, Gilad, Mirsky, Yisroel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
por: Rozenfeld, Shir, et al.
Publicado: (2026)
por: Rozenfeld, Shir, et al.
Publicado: (2026)
Who Owns This Agent? Tracing AI Agents Back to Their Owners
por: Chocron, Ruben, et al.
Publicado: (2026)
por: Chocron, Ruben, et al.
Publicado: (2026)
PEAS: A Strategy for Crafting Transferable Adversarial Examples
por: Avraham, Bar, et al.
Publicado: (2024)
por: Avraham, Bar, et al.
Publicado: (2024)
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
por: Bokobza, Roey, et al.
Publicado: (2024)
por: Bokobza, Roey, et al.
Publicado: (2024)
ProxyPrints: From Database Breach to Spoof, A Plug-and-Play Defense for Biometric Systems
por: Hacmon, Yaniv, et al.
Publicado: (2025)
por: Hacmon, Yaniv, et al.
Publicado: (2025)
The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks
por: Ayzenshteyn, Daniel, et al.
Publicado: (2024)
por: Ayzenshteyn, Daniel, et al.
Publicado: (2024)
Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
por: Gressel, Gilad, et al.
Publicado: (2025)
por: Gressel, Gilad, et al.
Publicado: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
por: Gressel, Gilad, et al.
Publicado: (2024)
por: Gressel, Gilad, et al.
Publicado: (2024)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
por: Bhatt, Manish, et al.
Publicado: (2026)
por: Bhatt, Manish, et al.
Publicado: (2026)
Transferability Ranking of Adversarial Examples
por: Levy, Mosh, et al.
Publicado: (2022)
por: Levy, Mosh, et al.
Publicado: (2022)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
por: Weiss, Roy, et al.
Publicado: (2024)
por: Weiss, Roy, et al.
Publicado: (2024)
Mapping LLM Security Landscapes: A Comprehensive Stakeholder Risk Assessment Proposal
por: Pankajakshan, Rahul, et al.
Publicado: (2024)
por: Pankajakshan, Rahul, et al.
Publicado: (2024)
Efficient Model Extraction via Boundary Sampling
por: Dor, Maor Biton, et al.
Publicado: (2024)
por: Dor, Maor Biton, et al.
Publicado: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
por: Halawi, Danny, et al.
Publicado: (2024)
por: Halawi, Danny, et al.
Publicado: (2024)
Real-time ML-based Defense Against Malicious Payload in Reconfigurable Embedded Systems
por: Stahle-Smith, Rye, et al.
Publicado: (2025)
por: Stahle-Smith, Rye, et al.
Publicado: (2025)
Shape and Substance: Dual-Layer Side-Channel Attacks on Local Vision-Language Models
por: Hadad, Eyal, et al.
Publicado: (2026)
por: Hadad, Eyal, et al.
Publicado: (2026)
ASRJam: Human-Friendly AI Speech Jamming to Prevent Automated Phone Scams
por: Grabovski, Freddie, et al.
Publicado: (2025)
por: Grabovski, Freddie, et al.
Publicado: (2025)
Transpose Attack: Stealing Datasets with Bidirectional Training
por: Amit, Guy, et al.
Publicado: (2023)
por: Amit, Guy, et al.
Publicado: (2023)
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
por: Shen, Xinjie, et al.
Publicado: (2026)
por: Shen, Xinjie, et al.
Publicado: (2026)
Attacks and Defenses Against LLM Fingerprinting
por: Kurian, Kevin, et al.
Publicado: (2025)
por: Kurian, Kevin, et al.
Publicado: (2025)
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
por: Guo, Yihao, et al.
Publicado: (2025)
por: Guo, Yihao, et al.
Publicado: (2025)
Optimal Defenses Against Gradient Reconstruction Attacks
por: Chen, Yuxiao, et al.
Publicado: (2024)
por: Chen, Yuxiao, et al.
Publicado: (2024)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
por: Rehman, Mohammad Abdul, et al.
Publicado: (2025)
por: Rehman, Mohammad Abdul, et al.
Publicado: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
por: Cao, Tri, et al.
Publicado: (2026)
por: Cao, Tri, et al.
Publicado: (2026)
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
por: Gloaguen, Thibaud, et al.
Publicado: (2025)
por: Gloaguen, Thibaud, et al.
Publicado: (2025)
Mitigating the Structural Bias in Graph Adversarial Defenses
por: Fang, Junyuan, et al.
Publicado: (2025)
por: Fang, Junyuan, et al.
Publicado: (2025)
SDD: Self-Degraded Defense against Malicious Fine-tuning
por: Chen, Zixuan, et al.
Publicado: (2025)
por: Chen, Zixuan, et al.
Publicado: (2025)
SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
por: Zhang, Zhisheng, et al.
Publicado: (2025)
por: Zhang, Zhisheng, et al.
Publicado: (2025)
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
por: Hasan, Md. Mehedi, et al.
Publicado: (2025)
por: Hasan, Md. Mehedi, et al.
Publicado: (2025)
Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
por: Thornton, Scott
Publicado: (2026)
por: Thornton, Scott
Publicado: (2026)
Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS
por: Ennaji, Sabrine, et al.
Publicado: (2025)
por: Ennaji, Sabrine, et al.
Publicado: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
por: Jiang, Weisen, et al.
Publicado: (2025)
por: Jiang, Weisen, et al.
Publicado: (2025)
TED-LaST: Towards Robust Backdoor Defense Against Adaptive Attacks
por: Mo, Xiaoxing, et al.
Publicado: (2025)
por: Mo, Xiaoxing, et al.
Publicado: (2025)
Adversarial Robustness in Financial Machine Learning: Defenses, Economic Impact, and Governance Evidence
por: Baviskar, Samruddhi
Publicado: (2025)
por: Baviskar, Samruddhi
Publicado: (2025)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
por: Campbell, David, et al.
Publicado: (2026)
por: Campbell, David, et al.
Publicado: (2026)
The Dark Side of Digital Twins: Adversarial Attacks on AI-Driven Water Forecasting
por: Homaei, Mohammadhossein, et al.
Publicado: (2025)
por: Homaei, Mohammadhossein, et al.
Publicado: (2025)
DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns
por: Phan, Trung V., et al.
Publicado: (2026)
por: Phan, Trung V., et al.
Publicado: (2026)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
por: Panebianco, Francesco, et al.
Publicado: (2025)
por: Panebianco, Francesco, et al.
Publicado: (2025)
AdaPhish: AI-Powered Adaptive Defense and Education Resource Against Deceptive Emails
por: Meguro, Rei, et al.
Publicado: (2025)
por: Meguro, Rei, et al.
Publicado: (2025)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
por: Casper, Stephen, et al.
Publicado: (2024)
por: Casper, Stephen, et al.
Publicado: (2024)
Ejemplares similares
-
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
por: Rozenfeld, Shir, et al.
Publicado: (2026) -
Who Owns This Agent? Tracing AI Agents Back to Their Owners
por: Chocron, Ruben, et al.
Publicado: (2026) -
PEAS: A Strategy for Crafting Transferable Adversarial Examples
por: Avraham, Bar, et al.
Publicado: (2024) -
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
por: Bokobza, Roey, et al.
Publicado: (2024) -
ProxyPrints: From Database Breach to Spoof, A Plug-and-Play Defense for Biometric Systems
por: Hacmon, Yaniv, et al.
Publicado: (2025)