Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Halawi, Danny, Wei, Alexander, Wallace, Eric, Wang, Tony T., Haghtalab, Nika, Steinhardt, Jacob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adversaries Can Misuse Combinations of Safe Models
von: Jones, Erik, et al.
Veröffentlicht: (2024)
von: Jones, Erik, et al.
Veröffentlicht: (2024)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
von: Zloczower, Itay, et al.
Veröffentlicht: (2026)
von: Zloczower, Itay, et al.
Veröffentlicht: (2026)
GuardReasoner: Towards Reasoning-based LLM Safeguards
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
von: Wahed, Muntasir, et al.
Veröffentlicht: (2025)
von: Wahed, Muntasir, et al.
Veröffentlicht: (2025)
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
von: Fonseca, Joao, et al.
Veröffentlicht: (2025)
von: Fonseca, Joao, et al.
Veröffentlicht: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
von: Yan, Bingyu, et al.
Veröffentlicht: (2026)
von: Yan, Bingyu, et al.
Veröffentlicht: (2026)
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
von: Wang, Kai, et al.
Veröffentlicht: (2026)
von: Wang, Kai, et al.
Veröffentlicht: (2026)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
von: Kaunismaa, Jackson, et al.
Veröffentlicht: (2026)
von: Kaunismaa, Jackson, et al.
Veröffentlicht: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
von: Wang, Tony T., et al.
Veröffentlicht: (2023)
von: Wang, Tony T., et al.
Veröffentlicht: (2023)
Approaching Human-Level Forecasting with Language Models
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Detecting Malicious AI Agents Through Simulated Interactions
von: Pi, Yulu, et al.
Veröffentlicht: (2025)
von: Pi, Yulu, et al.
Veröffentlicht: (2025)
Fair Finetuning Mitigates Distribution Inference Attacks
von: Naidu, Rakshit
Veröffentlicht: (2026)
von: Naidu, Rakshit
Veröffentlicht: (2026)
Finetuning Large Language Models for Vulnerability Detection
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
von: Shi, Zhan, et al.
Veröffentlicht: (2025)
von: Shi, Zhan, et al.
Veröffentlicht: (2025)
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
von: Liang, Zi, et al.
Veröffentlicht: (2025)
von: Liang, Zi, et al.
Veröffentlicht: (2025)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure Inference
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023)
von: Mishra, Abhijit, et al.
Veröffentlicht: (2023)
Federated In-Context LLM Agent Learning
von: Wu, Panlong, et al.
Veröffentlicht: (2024)
von: Wu, Panlong, et al.
Veröffentlicht: (2024)
Localizing Malicious Outputs from CodeLLM
von: Borana, Mayukh, et al.
Veröffentlicht: (2025)
von: Borana, Mayukh, et al.
Veröffentlicht: (2025)
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
von: Guo, Yihao, et al.
Veröffentlicht: (2025)
von: Guo, Yihao, et al.
Veröffentlicht: (2025)
Hierarchical Local-Global Feature Learning for Few-shot Malicious Traffic Detection
von: Peng, Songtao, et al.
Veröffentlicht: (2025)
von: Peng, Songtao, et al.
Veröffentlicht: (2025)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
von: Zhu, Sicheng, et al.
Veröffentlicht: (2024)
Policy-Invisible Violations in LLM-Based Agents
von: Wu, Jie, et al.
Veröffentlicht: (2026)
von: Wu, Jie, et al.
Veröffentlicht: (2026)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2025)
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2025)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
LLM Cyber Evaluations Don't Capture Real-World Risk
von: Lukošiūtė, Kamilė, et al.
Veröffentlicht: (2025)
von: Lukošiūtė, Kamilė, et al.
Veröffentlicht: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Adversaries Can Misuse Combinations of Safe Models
von: Jones, Erik, et al.
Veröffentlicht: (2024) -
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023) -
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
von: Zloczower, Itay, et al.
Veröffentlicht: (2026) -
GuardReasoner: Towards Reasoning-based LLM Safeguards
von: Liu, Yue, et al.
Veröffentlicht: (2025) -
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
von: Wahed, Muntasir, et al.
Veröffentlicht: (2025)