NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Haodong, Hu, Jinming, Li, Peixuan, Li, Fangqi, Sha, Jinrui, Ju, Tianjie, Chen, Peixuan, Zhang, Zhuosheng, Liu, Gongshen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EmbTracker: Traceable Black-box Watermarking for Federated Language Models
by: Zhao, Haodong, et al.
Published: (2026)
by: Zhao, Haodong, et al.
Published: (2026)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
by: Du, Wei, et al.
Published: (2023)
by: Du, Wei, et al.
Published: (2023)
Revisiting Backdoor Threat in Federated Instruction Tuning from a Signal Aggregation Perspective
by: Zhao, Haodong, et al.
Published: (2026)
by: Zhao, Haodong, et al.
Published: (2026)
Transferring Backdoors between Large Language Models by Knowledge Distillation
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
Revisiting the Information Capacity of Neural Network Watermarks: Upper Bound Estimation and Beyond
by: Li, Fangqi, et al.
Published: (2024)
by: Li, Fangqi, et al.
Published: (2024)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data
by: Zhao, Haodong, et al.
Published: (2026)
by: Zhao, Haodong, et al.
Published: (2026)
Performance-lossless Black-box Model Watermarking
by: Zhao, Na, et al.
Published: (2023)
by: Zhao, Na, et al.
Published: (2023)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
by: Wu, Zongru, et al.
Published: (2024)
by: Wu, Zongru, et al.
Published: (2024)
Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
by: Zhao, Tianhang, et al.
Published: (2025)
by: Zhao, Tianhang, et al.
Published: (2025)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
by: Zhu, Hongyu, et al.
Published: (2024)
by: Zhu, Hongyu, et al.
Published: (2024)
Traceable Black-box Watermarks for Federated Learning
by: Xu, Jiahao, et al.
Published: (2025)
by: Xu, Jiahao, et al.
Published: (2025)
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
by: Li, Hongyi, et al.
Published: (2024)
by: Li, Hongyi, et al.
Published: (2024)
Audio Pirates: Black-box Audio Watermark Removal via Diffusion Priors
by: Yao, Lingfeng, et al.
Published: (2026)
by: Yao, Lingfeng, et al.
Published: (2026)
Compositional security definitions for higher-order where declassification
by: Menz, Jan, et al.
Published: (2026)
by: Menz, Jan, et al.
Published: (2026)
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
by: Cheng, Pengzhou, et al.
Published: (2023)
by: Cheng, Pengzhou, et al.
Published: (2023)
LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
by: Yan, Zihe, et al.
Published: (2025)
by: Yan, Zihe, et al.
Published: (2025)
HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
Query Provenance Analysis: Efficient and Robust Defense against Query-based Black-box Attacks
by: Li, Shaofei, et al.
Published: (2024)
by: Li, Shaofei, et al.
Published: (2024)
Robust and Imperceptible Black-box DNN Watermarking Based on Fourier Perturbation Analysis and Frequency Sensitivity Clustering
by: Liu, Yong, et al.
Published: (2022)
by: Liu, Yong, et al.
Published: (2022)
SEW: Strengthening Robustness of Black-box DNN Watermarking via Specificity Enhancement
by: Qiu, Huming, et al.
Published: (2026)
by: Qiu, Huming, et al.
Published: (2026)
LoRAGuard: An Effective Black-box Watermarking Approach for LoRAs
by: Lv, Peizhuo, et al.
Published: (2025)
by: Lv, Peizhuo, et al.
Published: (2025)
Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models
by: Huang, Chung-ju, et al.
Published: (2026)
by: Huang, Chung-ju, et al.
Published: (2026)
A Game Between the Defender and the Attacker for Trigger-based Black-box Model Watermarking
by: Huang, Chaoyue, et al.
Published: (2025)
by: Huang, Chaoyue, et al.
Published: (2025)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
WGLE:Backdoor-free and Multi-bit Black-box Watermarking for Graph Neural Networks
by: Li, Tingzhi, et al.
Published: (2025)
by: Li, Tingzhi, et al.
Published: (2025)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
by: Wu, Zongru, et al.
Published: (2024)
by: Wu, Zongru, et al.
Published: (2024)
Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks
by: Xiao, Yaxin, et al.
Published: (2025)
by: Xiao, Yaxin, et al.
Published: (2025)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited Data
by: Lu, Yifan, et al.
Published: (2023)
by: Lu, Yifan, et al.
Published: (2023)
AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection
by: Gao, Jianbo, et al.
Published: (2025)
by: Gao, Jianbo, et al.
Published: (2025)
A Universal Identity Backdoor Attack against Speaker Verification based on Siamese Network
by: Zhao, Haodong, et al.
Published: (2023)
by: Zhao, Haodong, et al.
Published: (2023)
Bounding-box Watermarking: Defense against Model Extraction Attacks on Object Detectors
by: Koda, Satoru, et al.
Published: (2024)
by: Koda, Satoru, et al.
Published: (2024)
When There Is No Decoder: Removing Watermarks from Stable Diffusion Models in a No-box Setting
by: Wu, Xiaodong, et al.
Published: (2025)
by: Wu, Xiaodong, et al.
Published: (2025)
Lite-BD: A Lightweight Black-box Backdoor Defense via Reviving Multi-Stage Image Transformations
by: Miah, Abdullah Arafat, et al.
Published: (2026)
by: Miah, Abdullah Arafat, et al.
Published: (2026)
BDFirewall: Towards Effective and Expeditiously Black-Box Backdoor Defense in MLaaS
by: Li, Ye, et al.
Published: (2025)
by: Li, Ye, et al.
Published: (2025)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
by: Zhang, Zhuomeng, et al.
Published: (2024)
by: Zhang, Zhuomeng, et al.
Published: (2024)
Detecting Fraudulent Services on Quantum Cloud Platforms via Dynamic Fingerprinting
by: Wu, Jindi, et al.
Published: (2024)
by: Wu, Jindi, et al.
Published: (2024)
Revisiting Data Auditing in Large Vision-Language Models
by: Zhu, Hongyu, et al.
Published: (2025)
by: Zhu, Hongyu, et al.
Published: (2025)
Similar Items
-
EmbTracker: Traceable Black-box Watermarking for Federated Language Models
by: Zhao, Haodong, et al.
Published: (2026) -
UOR: Universal Backdoor Attacks on Pre-trained Language Models
by: Du, Wei, et al.
Published: (2023) -
Revisiting Backdoor Threat in Federated Instruction Tuning from a Signal Aggregation Perspective
by: Zhao, Haodong, et al.
Published: (2026) -
Transferring Backdoors between Large Language Models by Knowledge Distillation
by: Cheng, Pengzhou, et al.
Published: (2024) -
Revisiting the Information Capacity of Neural Network Watermarks: Upper Bound Estimation and Beyond
by: Li, Fangqi, et al.
Published: (2024)