Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
Fuente:
arXiv
Saved in:
| Main Authors: | Junhao, Wei, Zhe, Yu, Jun, Sakuma |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
by: Huang, Hanbo, et al.
Published: (2024)
by: Huang, Hanbo, et al.
Published: (2024)
DP-TRAE: A Dual-Phase Merging Transferable Reversible Adversarial Example for Image Privacy Protection
by: Du, Xia, et al.
Published: (2025)
by: Du, Xia, et al.
Published: (2025)
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
by: Tan, Yixin, et al.
Published: (2025)
by: Tan, Yixin, et al.
Published: (2025)
Differentially Private Model Merging
by: Yin, Qichuan, et al.
Published: (2026)
by: Yin, Qichuan, et al.
Published: (2026)
A Survey on Model Extraction Attacks and Defenses for Large Language Models
by: Zhao, Kaixiang, et al.
Published: (2025)
by: Zhao, Kaixiang, et al.
Published: (2025)
Attacks and Defenses for Generative Diffusion Models: A Comprehensive Survey
by: Truong, Vu Tuan, et al.
Published: (2024)
by: Truong, Vu Tuan, et al.
Published: (2024)
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
by: Leong, Jun Wen
Published: (2026)
by: Leong, Jun Wen
Published: (2026)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
A Survey of Model Extraction Attacks and Defenses in Distributed Computing Environments
by: Zhao, Kaixiang, et al.
Published: (2025)
by: Zhao, Kaixiang, et al.
Published: (2025)
LoBAM: LoRA-Based Backdoor Attack on Model Merging
by: Yin, Ming, et al.
Published: (2024)
by: Yin, Ming, et al.
Published: (2024)
A Survey of Privacy Threats and Defense in Vertical Federated Learning: From Model Life Cycle Perspective
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives
by: Zhao, Kaixiang, et al.
Published: (2025)
by: Zhao, Kaixiang, et al.
Published: (2025)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
by: Jaffal, Niveen O., et al.
Published: (2025)
by: Jaffal, Niveen O., et al.
Published: (2025)
A Factored MDP Approach To Moving Target Defense With Dynamic Threat Modeling and Cost Efficiency
by: Bose, Megha, et al.
Published: (2024)
by: Bose, Megha, et al.
Published: (2024)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
Attacks and Defenses Against LLM Fingerprinting
by: Kurian, Kevin, et al.
Published: (2025)
by: Kurian, Kevin, et al.
Published: (2025)
TBNet: A Neural Architectural Defense Framework Facilitating DNN Model Protection in Trusted Execution Environments
by: Liu, Ziyu, et al.
Published: (2024)
by: Liu, Ziyu, et al.
Published: (2024)
PM-MOE: Mixture of Experts on Private Model Parameters for Personalized Federated Learning
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
Privacy and Accuracy Implications of Model Complexity and Integration in Heterogeneous Federated Learning
by: Németh, Gergely Dániel, et al.
Published: (2023)
by: Németh, Gergely Dániel, et al.
Published: (2023)
Retrieval Augmented Anomaly Detection (RAAD): Nimble Model Adjustment Without Retraining
by: Pastoriza, Sam, et al.
Published: (2025)
by: Pastoriza, Sam, et al.
Published: (2025)
Mitigating the Structural Bias in Graph Adversarial Defenses
by: Fang, Junyuan, et al.
Published: (2025)
by: Fang, Junyuan, et al.
Published: (2025)
Optimal Defenses Against Gradient Reconstruction Attacks
by: Chen, Yuxiao, et al.
Published: (2024)
by: Chen, Yuxiao, et al.
Published: (2024)
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions
by: Li, Wenjuan, et al.
Published: (2026)
by: Li, Wenjuan, et al.
Published: (2026)
C2A: Client-Customized Adaptation for Parameter-Efficient Federated Learning
by: Kim, Yeachan, et al.
Published: (2024)
by: Kim, Yeachan, et al.
Published: (2024)
SHIELD: Secure Hypernetworks for Incremental Expansion Learning Defense
by: Krukowski, Patryk, et al.
Published: (2025)
by: Krukowski, Patryk, et al.
Published: (2025)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
The Autonomy Tax: Defense Training Breaks LLM Agents
by: Li, Shawn, et al.
Published: (2026)
by: Li, Shawn, et al.
Published: (2026)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
PureGen: Universal Data Purification for Train-Time Poison Defense via Generative Model Dynamics
by: Bhat, Sunay, et al.
Published: (2024)
by: Bhat, Sunay, et al.
Published: (2024)
Training RL Agents for Multi-Objective Network Defense Tasks
by: Molina-Markham, Andres, et al.
Published: (2025)
by: Molina-Markham, Andres, et al.
Published: (2025)
Seed Hijacking of LLM Sampling and Quantum Random Number Defense
by: You, Ziyang, et al.
Published: (2026)
by: You, Ziyang, et al.
Published: (2026)
Combining Stochastic Defenses to Resist Gradient Inversion: An Ablation Study
by: Scheliga, Daniel, et al.
Published: (2022)
by: Scheliga, Daniel, et al.
Published: (2022)
PoolFlip: A Multi-Agent Reinforcement Learning Security Environment for Cyber Defense
by: Cadet, Xavier, et al.
Published: (2025)
by: Cadet, Xavier, et al.
Published: (2025)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
by: Panebianco, Francesco, et al.
Published: (2025)
by: Panebianco, Francesco, et al.
Published: (2025)
Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses
by: Yichao, Wu, et al.
Published: (2025)
by: Yichao, Wu, et al.
Published: (2025)
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
by: Konrad, Phongsakon Mark, et al.
Published: (2026)
by: Konrad, Phongsakon Mark, et al.
Published: (2026)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
Optimizing Cyber Defense in Dynamic Active Directories through Reinforcement Learning
by: Goel, Diksha, et al.
Published: (2024)
by: Goel, Diksha, et al.
Published: (2024)
Similar Items
-
A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
by: Huang, Hanbo, et al.
Published: (2024) -
DP-TRAE: A Dual-Phase Merging Transferable Reversible Adversarial Example for Image Privacy Protection
by: Du, Xia, et al.
Published: (2025) -
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
by: Tan, Yixin, et al.
Published: (2025) -
Differentially Private Model Merging
by: Yin, Qichuan, et al.
Published: (2026) -
A Survey on Model Extraction Attacks and Defenses for Large Language Models
by: Zhao, Kaixiang, et al.
Published: (2025)