On the Weaknesses of Backdoor-based Model Watermarking: An Information-theoretic Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Aoting, Chen, Yanzhi, Xie, Renjie, Weller, Adrian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Coward: Collision-based OOD Watermarking for Practical Proactive Federated Backdoor Detection
von: Li, Wenjie, et al.
Veröffentlicht: (2025)
von: Li, Wenjie, et al.
Veröffentlicht: (2025)
Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model Watermarking
von: Kong, Cong, et al.
Veröffentlicht: (2024)
von: Kong, Cong, et al.
Veröffentlicht: (2024)
WGLE:Backdoor-free and Multi-bit Black-box Watermarking for Graph Neural Networks
von: Li, Tingzhi, et al.
Veröffentlicht: (2025)
von: Li, Tingzhi, et al.
Veröffentlicht: (2025)
SSCL-BW: Sample-Specific Clean-Label Backdoor Watermarking for Dataset Ownership Verification
von: Wang, Yingjia, et al.
Veröffentlicht: (2025)
von: Wang, Yingjia, et al.
Veröffentlicht: (2025)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
Investigating Deep Watermark Security: An Adversarial Transferability Perspective
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
Functional Subspace Watermarking for Large Language Models
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2025)
von: Yu, Miao, et al.
Veröffentlicht: (2025)
Large Language Model Watermark Stealing With Mixed Integer Programming
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024)
Invariant-based Robust Weights Watermark for Large Language Models
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
Selection-Based Vulnerabilities: Clean-Label Backdoor Attacks in Active Learning
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
von: Hu, Man, et al.
Veröffentlicht: (2025)
von: Hu, Man, et al.
Veröffentlicht: (2025)
Inevitable Trade-off between Watermark Strength and Speculative Sampling Efficiency for Language Models
von: Hu, Zhengmian, et al.
Veröffentlicht: (2024)
von: Hu, Zhengmian, et al.
Veröffentlicht: (2024)
SFIBA: Spatial-based Full-target Invisible Backdoor Attacks
von: Yin, Yangxu, et al.
Veröffentlicht: (2025)
von: Yin, Yangxu, et al.
Veröffentlicht: (2025)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
von: Zhu, Hongyu, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyu, et al.
Veröffentlicht: (2024)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency
von: Wang, Bingzheng, et al.
Veröffentlicht: (2026)
von: Wang, Bingzheng, et al.
Veröffentlicht: (2026)
A Survey of Fragile Model Watermarking
von: Gao, Zhenzhe, et al.
Veröffentlicht: (2024)
von: Gao, Zhenzhe, et al.
Veröffentlicht: (2024)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
Backdooring Bias in Large Language Models
von: Das, Anudeep, et al.
Veröffentlicht: (2026)
von: Das, Anudeep, et al.
Veröffentlicht: (2026)
Lightweight and Fast Backdoor Model Detection
von: Yu, Yinbo, et al.
Veröffentlicht: (2026)
von: Yu, Yinbo, et al.
Veröffentlicht: (2026)
Modification and Generated-Text Detection: Achieving Dual Detection Capabilities for the Outputs of LLM by Watermark
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)
FFCBA: Feature-based Full-target Clean-label Backdoor Attacks
von: Yin, Yangxu, et al.
Veröffentlicht: (2025)
von: Yin, Yangxu, et al.
Veröffentlicht: (2025)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
von: Ma, Yuan, et al.
Veröffentlicht: (2024)
von: Ma, Yuan, et al.
Veröffentlicht: (2024)
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
von: Popovic, Dorde, et al.
Veröffentlicht: (2025)
von: Popovic, Dorde, et al.
Veröffentlicht: (2025)
Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems
von: Jiao, Ruochen, et al.
Veröffentlicht: (2024)
von: Jiao, Ruochen, et al.
Veröffentlicht: (2024)
Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA Models
von: Chen, Linzhi, et al.
Veröffentlicht: (2025)
von: Chen, Linzhi, et al.
Veröffentlicht: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
von: Li, Xi, et al.
Veröffentlicht: (2024)
von: Li, Xi, et al.
Veröffentlicht: (2024)
Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
von: Wang, Chenrui, et al.
Veröffentlicht: (2025)
von: Wang, Chenrui, et al.
Veröffentlicht: (2025)
Invisible Textual Backdoor Attacks based on Dual-Trigger
von: Hou, Yang, et al.
Veröffentlicht: (2024)
von: Hou, Yang, et al.
Veröffentlicht: (2024)
BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
Multi-Designated Detector Watermarking for Language Models
von: Huang, Zhengan, et al.
Veröffentlicht: (2024)
von: Huang, Zhengan, et al.
Veröffentlicht: (2024)
The Coding Limits of Robust Watermarking for Generative Models
von: Francati, Danilo, et al.
Veröffentlicht: (2025)
von: Francati, Danilo, et al.
Veröffentlicht: (2025)
On Protecting Agentic Systems' Intellectual Property via Watermarking
von: Wang, Liwen, et al.
Veröffentlicht: (2026)
von: Wang, Liwen, et al.
Veröffentlicht: (2026)
Towards Backdoor Stealthiness in Model Parameter Space
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
CEFW: A Comprehensive Evaluation Framework for Watermark in Large Language Models
von: Zhang, Shuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Shuhao, et al.
Veröffentlicht: (2025)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
von: Li, Yige, et al.
Veröffentlicht: (2026)
von: Li, Yige, et al.
Veröffentlicht: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
von: Li, Yige, et al.
Veröffentlicht: (2025)
von: Li, Yige, et al.
Veröffentlicht: (2025)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Coward: Collision-based OOD Watermarking for Practical Proactive Federated Backdoor Detection
von: Li, Wenjie, et al.
Veröffentlicht: (2025) -
Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model Watermarking
von: Kong, Cong, et al.
Veröffentlicht: (2024) -
WGLE:Backdoor-free and Multi-bit Black-box Watermarking for Graph Neural Networks
von: Li, Tingzhi, et al.
Veröffentlicht: (2025) -
SSCL-BW: Sample-Specific Clean-Label Backdoor Watermarking for Dataset Ownership Verification
von: Wang, Yingjia, et al.
Veröffentlicht: (2025) -
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)