Quantized Delta Weight Is Safety Keeper
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yule, Sun, Zhen, He, Xinlei, Huang, Xinyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
Privacy-Preserving Federated Learning via Homomorphic Adversarial Networks
by: Dong, Wenhan, et al.
Published: (2024)
by: Dong, Wenhan, et al.
Published: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025)
by: Zheng, Jingyi, et al.
Published: (2025)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
by: Lu, Ning, et al.
Published: (2025)
by: Lu, Ning, et al.
Published: (2025)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
Verification of Bit-Flip Attacks against Quantized Neural Networks
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Mind the Gap: A Practical Attack on GGUF Quantization
by: Egashira, Kazuki, et al.
Published: (2025)
by: Egashira, Kazuki, et al.
Published: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
PromptKeeper: Safeguarding System Prompts for LLMs
by: Jiang, Zhifeng, et al.
Published: (2024)
by: Jiang, Zhifeng, et al.
Published: (2024)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
by: Chen, Xiangxiang, et al.
Published: (2025)
by: Chen, Xiangxiang, et al.
Published: (2025)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
by: Liu, Zesen, et al.
Published: (2024)
by: Liu, Zesen, et al.
Published: (2024)
Saffron-1: Safety Inference Scaling
by: Qiu, Ruizhong, et al.
Published: (2025)
by: Qiu, Ruizhong, et al.
Published: (2025)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Investigating the Impact of Quantization on Adversarial Robustness
by: Li, Qun, et al.
Published: (2024)
by: Li, Qun, et al.
Published: (2024)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
UpSafe$^\circ$C: Upcycling for Controllable Safety in Large Language Models
by: Sun, Yuhao, et al.
Published: (2025)
by: Sun, Yuhao, et al.
Published: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Aggressive Compression Enables LLM Weight Theft
by: Brown, Davis, et al.
Published: (2026)
by: Brown, Davis, et al.
Published: (2026)
Phishing Detection in the Gen-AI Era: Quantized LLMs vs Classical Models
by: Thapa, Jikesh, et al.
Published: (2025)
by: Thapa, Jikesh, et al.
Published: (2025)
PACZero: PAC-Private Fine-Tuning of Language Models via Sign Quantization
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols
by: Maheri, Mohammad M, et al.
Published: (2025)
by: Maheri, Mohammad M, et al.
Published: (2025)
RQP-SGD: Differential Private Machine Learning through Noisy SGD and Randomized Quantization
by: Feng, Ce, et al.
Published: (2024)
by: Feng, Ce, et al.
Published: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Self-Mined Hardness for Safety Fine-Tuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
by: Fang, Xingli, et al.
Published: (2026)
by: Fang, Xingli, et al.
Published: (2026)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
Why Safety Probes Catch Liars But Miss Fanatics
by: Haralambiev, Kristiyan
Published: (2026)
by: Haralambiev, Kristiyan
Published: (2026)
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
by: Egashira, Kazuki, et al.
Published: (2025)
by: Egashira, Kazuki, et al.
Published: (2025)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
A Safety and Security Framework for Real-World Agentic Systems
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
by: Rozenfeld, Shir, et al.
Published: (2026)
by: Rozenfeld, Shir, et al.
Published: (2026)
Similar Items
-
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025) -
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025) -
Privacy-Preserving Federated Learning via Homomorphic Adversarial Networks
by: Dong, Wenhan, et al.
Published: (2024) -
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024) -
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
by: Zheng, Jingyi, et al.
Published: (2025)