Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
Fuente:
arXiv
Saved in:
| Main Authors: | Gloaguen, Thibaud, Vero, Mark, Staab, Robin, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
by: Egashira, Kazuki, et al.
Published: (2025)
by: Egashira, Kazuki, et al.
Published: (2025)
Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
A Unified Framework for LLM Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
Watermarking Diffusion Language Models
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
Discovering Spoofing Attempts on Language Model Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2024)
by: Gloaguen, Thibaud, et al.
Published: (2024)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
LLM Fingerprinting via Semantically Conditioned Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
by: De Muri, Giovanni, et al.
Published: (2025)
by: De Muri, Giovanni, et al.
Published: (2025)
Mind the Gap: A Practical Attack on GGUF Quantization
by: Egashira, Kazuki, et al.
Published: (2025)
by: Egashira, Kazuki, et al.
Published: (2025)
Towards Watermarking of Open-Source LLMs
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
Black-Box Detection of Language Model Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2024)
by: Gloaguen, Thibaud, et al.
Published: (2024)
Ward: Provable RAG Dataset Inference via LLM Watermarks
by: Jovanović, Nikola, et al.
Published: (2024)
by: Jovanović, Nikola, et al.
Published: (2024)
Watermark Stealing in Large Language Models
by: Jovanović, Nikola, et al.
Published: (2024)
by: Jovanović, Nikola, et al.
Published: (2024)
Large Language Models are Advanced Anonymizers
by: Staab, Robin, et al.
Published: (2024)
by: Staab, Robin, et al.
Published: (2024)
Instruction Tuning for Secure Code Generation
by: He, Jingxuan, et al.
Published: (2024)
by: He, Jingxuan, et al.
Published: (2024)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
by: Vero, Mark, et al.
Published: (2026)
by: Vero, Mark, et al.
Published: (2026)
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
by: von Arx, Tobias, et al.
Published: (2025)
by: von Arx, Tobias, et al.
Published: (2025)
Black-Box Adversarial Attacks on LLM-Based Code Completion
by: Jenko, Slobodan, et al.
Published: (2024)
by: Jenko, Slobodan, et al.
Published: (2024)
BaxBench: Can LLMs Generate Correct and Secure Backends?
by: Vero, Mark, et al.
Published: (2025)
by: Vero, Mark, et al.
Published: (2025)
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
by: Zhan, Xiaohua, et al.
Published: (2026)
by: Zhan, Xiaohua, et al.
Published: (2026)
SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
A Synthetic Dataset for Personal Attribute Inference
by: Yukhymenko, Hanna, et al.
Published: (2024)
by: Yukhymenko, Hanna, et al.
Published: (2024)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)
by: Zloczower, Itay, et al.
Published: (2026)
Beyond Memorization: Violating Privacy Via Inference with Large Language Models
by: Staab, Robin, et al.
Published: (2023)
by: Staab, Robin, et al.
Published: (2023)
Private Attribute Inference from Images with Vision-Language Models
by: Tömekçe, Batuhan, et al.
Published: (2024)
by: Tömekçe, Batuhan, et al.
Published: (2024)
Back to the Drawing Board for Fair Representation Learning
by: Pouget, Angéline, et al.
Published: (2024)
by: Pouget, Angéline, et al.
Published: (2024)
Black-box Adversarial Attacks on Network-wide Multi-step Traffic State Prediction Models
by: Poudel, Bibek, et al.
Published: (2021)
by: Poudel, Bibek, et al.
Published: (2021)
Are Robust LLM Fingerprints Adversarially Robust?
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
Fair Finetuning Mitigates Distribution Inference Attacks
by: Naidu, Rakshit
Published: (2026)
by: Naidu, Rakshit
Published: (2026)
Finetuning Large Language Models for Vulnerability Detection
by: Shestov, Alexey, et al.
Published: (2024)
by: Shestov, Alexey, et al.
Published: (2024)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models
by: Raza, Ali, et al.
Published: (2026)
by: Raza, Ali, et al.
Published: (2026)
Modeling Behavioral Preferences of Cyber Adversaries Using Inverse Reinforcement Learning
by: Shinde, Aditya, et al.
Published: (2025)
by: Shinde, Aditya, et al.
Published: (2025)
SoK: Data Minimization in Machine Learning
by: Staab, Robin, et al.
Published: (2025)
by: Staab, Robin, et al.
Published: (2025)
Have it your way: Individualized Privacy Assignment for DP-SGD
by: Boenisch, Franziska, et al.
Published: (2023)
by: Boenisch, Franziska, et al.
Published: (2023)
Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
by: Chia, Xin Wei, et al.
Published: (2025)
by: Chia, Xin Wei, et al.
Published: (2025)
RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors
by: Bai, Fengshuo, et al.
Published: (2024)
by: Bai, Fengshuo, et al.
Published: (2024)
Vision Transformer with Adversarial Indicator Token against Adversarial Attacks in Radio Signal Classifications
by: Zhang, Lu, et al.
Published: (2025)
by: Zhang, Lu, et al.
Published: (2025)
Similar Items
-
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
by: Egashira, Kazuki, et al.
Published: (2025) -
Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark
by: Gloaguen, Thibaud, et al.
Published: (2026) -
A Unified Framework for LLM Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2026) -
Watermarking Diffusion Language Models
by: Gloaguen, Thibaud, et al.
Published: (2025) -
Discovering Spoofing Attempts on Language Model Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2024)