DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | An, Hyeseon, Park, Shinwoo, Woo, Suyeon, Han, Yo-Sub |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sequential Behavioral Watermarking for LLM Agents
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
by: Kim, Jungin, et al.
Published: (2025)
by: Kim, Jungin, et al.
Published: (2025)
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
Linguistics-Aware Non-Distortionary LLM Watermarking
by: Park, Shinwoo, et al.
Published: (2026)
by: Park, Shinwoo, et al.
Published: (2026)
WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
by: Shen, Huanming, et al.
Published: (2025)
by: Shen, Huanming, et al.
Published: (2025)
Steering Language Models Before They Speak: Logit-Level Interventions
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
Discovering Spoofing Attempts on Language Model Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2024)
by: Gloaguen, Thibaud, et al.
Published: (2024)
Experimental Validation of Sensor Fusion-based GNSS Spoofing Attack Detection Framework for Autonomous Vehicles
by: Dasgupta, Sagar, et al.
Published: (2024)
by: Dasgupta, Sagar, et al.
Published: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
by: Xu, Yixiao, et al.
Published: (2025)
by: Xu, Yixiao, et al.
Published: (2025)
LLM Watermark Evasion via Bias Inversion
by: Hwang, Jeongyeon, et al.
Published: (2025)
by: Hwang, Jeongyeon, et al.
Published: (2025)
Robust Knowledge Distillation in Federated Learning: Counteracting Backdoor Attacks
by: Alharbi, Ebtisaam, et al.
Published: (2025)
by: Alharbi, Ebtisaam, et al.
Published: (2025)
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
by: Fei, Zekun, et al.
Published: (2024)
by: Fei, Zekun, et al.
Published: (2024)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025)
by: Fu, Haowei, et al.
Published: (2025)
Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
by: Wang, Chenrui, et al.
Published: (2025)
by: Wang, Chenrui, et al.
Published: (2025)
Watermark Overwriting Attack on StegaStamp algorithm
by: Serzhenko, I. F., et al.
Published: (2025)
by: Serzhenko, I. F., et al.
Published: (2025)
FTSmartAudit: A Knowledge Distillation-Enhanced Framework for Automated Smart Contract Auditing Using Fine-Tuned LLMs
by: Wei, Zhiyuan, et al.
Published: (2024)
by: Wei, Zhiyuan, et al.
Published: (2024)
RTLMarker: Protecting LLM-Generated RTL Copyright via a Hardware Watermarking Framework
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
by: Liu, Qingyu, et al.
Published: (2026)
by: Liu, Qingyu, et al.
Published: (2026)
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024)
by: Diaa, Abdulrahman, et al.
Published: (2024)
On Membership Inference Attacks in Knowledge Distillation
by: Cui, Ziyao, et al.
Published: (2025)
by: Cui, Ziyao, et al.
Published: (2025)
CEFW: A Comprehensive Evaluation Framework for Watermark in Large Language Models
by: Zhang, Shuhao, et al.
Published: (2025)
by: Zhang, Shuhao, et al.
Published: (2025)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
From Intuition to Calibrated Judgment: A Rubric-Based Expert-Panel Study of Human Detection of LLM-Generated Korean Text
by: Park, Shinwoo, et al.
Published: (2026)
by: Park, Shinwoo, et al.
Published: (2026)
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
by: You, Ziyang, et al.
Published: (2026)
by: You, Ziyang, et al.
Published: (2026)
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025)
by: Cui, Jing, et al.
Published: (2025)
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
by: Yueh-Han, Chen, et al.
Published: (2025)
by: Yueh-Han, Chen, et al.
Published: (2025)
GPS Spoofing Attack Detection in Autonomous Vehicles Using Adaptive DBSCAN
by: Mohammadi, Ahmad, et al.
Published: (2025)
by: Mohammadi, Ahmad, et al.
Published: (2025)
Supervised and Unsupervised Alignments for Spoofing Behavioral Biometrics
by: Thebaud, Thomas, et al.
Published: (2024)
by: Thebaud, Thomas, et al.
Published: (2024)
PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
by: Ai, Zhenxin, et al.
Published: (2026)
by: Ai, Zhenxin, et al.
Published: (2026)
Clean-Label Physical Backdoor Attacks with Data Distillation
by: Dao, Thinh, et al.
Published: (2024)
by: Dao, Thinh, et al.
Published: (2024)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
FedSecurity: Benchmarking Attacks and Defenses in Federated Learning and Federated LLMs
by: Han, Shanshan, et al.
Published: (2023)
by: Han, Shanshan, et al.
Published: (2023)
On Protecting Agentic Systems' Intellectual Property via Watermarking
by: Wang, Liwen, et al.
Published: (2026)
by: Wang, Liwen, et al.
Published: (2026)
Similar Items
-
Sequential Behavioral Watermarking for LLM Agents
by: An, Hyeseon, et al.
Published: (2026) -
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
by: Kim, Jungin, et al.
Published: (2025) -
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
by: Park, Shinwoo, et al.
Published: (2025) -
Linguistics-Aware Non-Distortionary LLM Watermarking
by: Park, Shinwoo, et al.
Published: (2026) -
WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
by: Park, Shinwoo, et al.
Published: (2025)