Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Quoc Minh, Le, Trung, Wu, Jing, Bui, Anh Tuan, Harandi, Mehrtash |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Erasing Undesirable Influence in Diffusion Models
by: Wu, Jing, et al.
Published: (2024)
by: Wu, Jing, et al.
Published: (2024)
Scissorhands: Scrub Data Influence via Connection Sensitivity in Networks
by: Wu, Jing, et al.
Published: (2024)
by: Wu, Jing, et al.
Published: (2024)
Optimizing Specific and Shared Parameters for Efficient Parameter Tuning
by: Nguyen, Van-Anh, et al.
Published: (2025)
by: Nguyen, Van-Anh, et al.
Published: (2025)
Large-Scale Data-Free Knowledge Distillation for ImageNet via Multi-Resolution Data Generation
by: Tran, Minh-Tuan, et al.
Published: (2024)
by: Tran, Minh-Tuan, et al.
Published: (2024)
Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise
by: Luo, Haocheng, et al.
Published: (2025)
by: Luo, Haocheng, et al.
Published: (2025)
MUNBa: Machine Unlearning via Nash Bargaining
by: Wu, Jing, et al.
Published: (2024)
by: Wu, Jing, et al.
Published: (2024)
Text-Enhanced Data-free Approach for Federated Class-Incremental Learning
by: Tran, Minh-Tuan, et al.
Published: (2024)
by: Tran, Minh-Tuan, et al.
Published: (2024)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge Distillation
by: Tran, Minh-Tuan, et al.
Published: (2023)
by: Tran, Minh-Tuan, et al.
Published: (2023)
Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modeling
by: Tran, Minh-Tuan, et al.
Published: (2026)
by: Tran, Minh-Tuan, et al.
Published: (2026)
NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning
by: Yi, Xin, et al.
Published: (2024)
by: Yi, Xin, et al.
Published: (2024)
Explicit Eigenvalue Regularization Improves Sharpness-Aware Minimization
by: Luo, Haocheng, et al.
Published: (2025)
by: Luo, Haocheng, et al.
Published: (2025)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
by: Trung, Bui The, et al.
Published: (2026)
by: Trung, Bui The, et al.
Published: (2026)
Provably Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function
by: Le, Tung Quoc, et al.
Published: (2026)
by: Le, Tung Quoc, et al.
Published: (2026)
Exemplar-Free Continual Learning for State Space Models
by: Lee, Isaac Ning, et al.
Published: (2025)
by: Lee, Isaac Ning, et al.
Published: (2025)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Improving Generalization in Heterogeneous Federated Continual Learning via Spatio-Temporal Gradient Matching with Prototypical Coreset
by: Nguyen, Minh-Duong, et al.
Published: (2025)
by: Nguyen, Minh-Duong, et al.
Published: (2025)
Sharpness-Aware Teleportation on Riemannian Manifolds
by: Truong, Tuan, et al.
Published: (2023)
by: Truong, Tuan, et al.
Published: (2023)
Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink
by: Liu, Guozhi, et al.
Published: (2026)
by: Liu, Guozhi, et al.
Published: (2026)
Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge
by: Lu, Weikai, et al.
Published: (2024)
by: Lu, Weikai, et al.
Published: (2024)
XMainframe: A Large Language Model for Mainframe Modernization
by: Dau, Anh T. V., et al.
Published: (2024)
by: Dau, Anh T. V., et al.
Published: (2024)
IMU: Influence-guided Machine Unlearning
by: Fan, Xindi, et al.
Published: (2025)
by: Fan, Xindi, et al.
Published: (2025)
Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
Modeling‐Facilitated Field Survey Discovers of a New Population of the Annamite Striped Rabbit in Kon Tum Province, Vietnam
by: Anh Tuan Nguyen, et al.
Published: (2024)
by: Anh Tuan Nguyen, et al.
Published: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
Provably Data-driven Lagrangian Relaxation for Mixed Integer Linear Programming
by: Le, Tung Quoc, et al.
Published: (2026)
by: Le, Tung Quoc, et al.
Published: (2026)
Multivariate Statistical Analysis for the Classification of Sausages Based on Physicochemical Attributes, Using Attenuated Total Reflectance-Fourier Transform Infrared (ATR-FTIR) and Inductively Coupled Plasma-Mass Spectrometry (ICP-MS)
by: Quang Minh Bui, et al.
Published: (2024)
by: Quang Minh Bui, et al.
Published: (2024)
Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme Detection
by: Pan, Fengjun, et al.
Published: (2025)
by: Pan, Fengjun, et al.
Published: (2025)
Low-energy $^7$Li($n,γ$)$^8$Li and $^7$Be($p,γ$)$^8$B radiative capture reactions within the Skyrme Hartree-Fock approach
by: Nguyen, Le-Anh, et al.
Published: (2022)
by: Nguyen, Le-Anh, et al.
Published: (2022)
Direct correlation between the near-proton-emission threshold resonance in $^{11}$B and the branching ratio of beta-delayed proton emission from $^{11}$Be
by: Nguyen, Le-Anh, et al.
Published: (2024)
by: Nguyen, Le-Anh, et al.
Published: (2024)
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025)
by: Kim, Heehwan, et al.
Published: (2025)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2026)
by: Van Vo, Tuan, et al.
Published: (2026)
Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation
by: Liu, Guozhi, et al.
Published: (2024)
by: Liu, Guozhi, et al.
Published: (2024)
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
by: Le, Nguyen-Khang, et al.
Published: (2025)
by: Le, Nguyen-Khang, et al.
Published: (2025)
Why Domain Generalization Fail? A View of Necessity and Sufficiency
by: Vuong, Long-Tung, et al.
Published: (2025)
by: Vuong, Long-Tung, et al.
Published: (2025)
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2025)
by: Van Vo, Tuan, et al.
Published: (2025)
PADM: A Physics-aware Diffusion Model for Attenuation Correction
by: Pham, Trung Kien, et al.
Published: (2025)
by: Pham, Trung Kien, et al.
Published: (2025)
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
by: Luo, Haocheng, et al.
Published: (2026)
by: Luo, Haocheng, et al.
Published: (2026)
Similar Items
-
Erasing Undesirable Influence in Diffusion Models
by: Wu, Jing, et al.
Published: (2024) -
Scissorhands: Scrub Data Influence via Connection Sensitivity in Networks
by: Wu, Jing, et al.
Published: (2024) -
Optimizing Specific and Shared Parameters for Efficient Parameter Tuning
by: Nguyen, Van-Anh, et al.
Published: (2025) -
Large-Scale Data-Free Knowledge Distillation for ImageNet via Multi-Resolution Data Generation
by: Tran, Minh-Tuan, et al.
Published: (2024) -
Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise
by: Luo, Haocheng, et al.
Published: (2025)