Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?
Fuente:
arXiv
Guardado en:
| Autores principales: | Panaitescu-Liess, Michael-Andrei, Che, Zora, An, Bang, Xu, Yuancheng, Pathmanathan, Pankayaraj, Chakraborty, Souradip, Zhu, Sicheng, Goldstein, Tom, Huang, Furong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
por: Panaitescu-Liess, Michael-Andrei, et al.
Publicado: (2025)
por: Panaitescu-Liess, Michael-Andrei, et al.
Publicado: (2025)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024)
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024)
RAGPart & RAGMask: Retrieval-Stage Defenses Against Corpus Poisoning in Retrieval-Augmented Generation
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2025)
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2025)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
por: An, Bang, et al.
Publicado: (2024)
por: An, Bang, et al.
Publicado: (2024)
Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2025)
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2025)
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
por: An, Bang, et al.
Publicado: (2023)
por: An, Bang, et al.
Publicado: (2023)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024)
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024)
Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2026)
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2026)
Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
por: Panaitescu-Liess, Michael-Andrei, et al.
Publicado: (2025)
por: Panaitescu-Liess, Michael-Andrei, et al.
Publicado: (2025)
Compositional Adversarial Training for Robust Visual Watermarking
por: Satheesh, Anirudh, et al.
Publicado: (2026)
por: Satheesh, Anirudh, et al.
Publicado: (2026)
WAVES: Benchmarking the Robustness of Image Watermarks
por: An, Bang, et al.
Publicado: (2024)
por: An, Bang, et al.
Publicado: (2024)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
por: Ding, Mucong, et al.
Publicado: (2024)
por: Ding, Mucong, et al.
Publicado: (2024)
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
por: Cai, Zikui, et al.
Publicado: (2025)
por: Cai, Zikui, et al.
Publicado: (2025)
Using Curiosity for an Even Representation of Tasks in Continual Offline Reinforcement Learning
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2023)
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2023)
Watermarking Text Data on Large Language Models for Dataset Copyright
por: Liu, Yixin, et al.
Publicado: (2023)
por: Liu, Yixin, et al.
Publicado: (2023)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
por: Xu, Yuancheng, et al.
Publicado: (2024)
por: Xu, Yuancheng, et al.
Publicado: (2024)
Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL
por: Liu, Xiangyu, et al.
Publicado: (2023)
por: Liu, Xiangyu, et al.
Publicado: (2023)
Agentic Critical Training
por: Liu, Weize, et al.
Publicado: (2026)
por: Liu, Weize, et al.
Publicado: (2026)
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
por: Agrawal, Aakriti, et al.
Publicado: (2024)
por: Agrawal, Aakriti, et al.
Publicado: (2024)
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
por: Agrawal, Aakriti, et al.
Publicado: (2025)
por: Agrawal, Aakriti, et al.
Publicado: (2025)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
por: Xu, Yuancheng, et al.
Publicado: (2024)
por: Xu, Yuancheng, et al.
Publicado: (2024)
SAFLEX: Self-Adaptive Augmentation via Feature Label Extrapolation
por: Ding, Mucong, et al.
Publicado: (2024)
por: Ding, Mucong, et al.
Publicado: (2024)
GATES: Self-Distillation under Privileged Context with Consensus Gating
por: Stein, Alex, et al.
Publicado: (2026)
por: Stein, Alex, et al.
Publicado: (2026)
A Watermark for Large Language Models
por: Kirchenbauer, John, et al.
Publicado: (2023)
por: Kirchenbauer, John, et al.
Publicado: (2023)
We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice
por: Petrov, Aleksandar, et al.
Publicado: (2025)
por: Petrov, Aleksandar, et al.
Publicado: (2025)
Analysis of Two Models for the Angular Structure of the Outflows Producing the Swift/XRT "Larger-Angle Emission" of Gamma-Ray Bursts
por: Panaitescu, A.
Publicado: (2025)
por: Panaitescu, A.
Publicado: (2025)
Can ChatGPT Perform Image Splicing Detection? A Preliminary Study
por: Nath, Souradip
Publicado: (2025)
por: Nath, Souradip
Publicado: (2025)
Can Protective Watermarking Safeguard the Copyright of 3D Gaussian Splatting?
por: Huang, Wenkai, et al.
Publicado: (2025)
por: Huang, Wenkai, et al.
Publicado: (2025)
Are Large Random Graphs Always Safe to Hide?
por: Chakraborty, Sourav, et al.
Publicado: (2025)
por: Chakraborty, Sourav, et al.
Publicado: (2025)
DataSafe: Copyright Protection with PUF Watermarking and Blockchain Tracking
por: Xue, Xiaolong, et al.
Publicado: (2024)
por: Xue, Xiaolong, et al.
Publicado: (2024)
Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models
por: He, Zhiwei, et al.
Publicado: (2024)
por: He, Zhiwei, et al.
Publicado: (2024)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
por: Agrawal, Aakriti, et al.
Publicado: (2025)
por: Agrawal, Aakriti, et al.
Publicado: (2025)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
por: Beetham, James, et al.
Publicado: (2024)
por: Beetham, James, et al.
Publicado: (2024)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
por: Barakat, Anas, et al.
Publicado: (2026)
por: Barakat, Anas, et al.
Publicado: (2026)
Auction-Based Regulation for Artificial Intelligence
por: Bornstein, Marco, et al.
Publicado: (2024)
por: Bornstein, Marco, et al.
Publicado: (2024)
On the Reliability of Watermarks for Large Language Models
por: Kirchenbauer, John, et al.
Publicado: (2023)
por: Kirchenbauer, John, et al.
Publicado: (2023)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
por: Pan, Leyi, et al.
Publicado: (2025)
por: Pan, Leyi, et al.
Publicado: (2025)
Watermark-embedded Adversarial Examples for Copyright Protection against Diffusion Models
por: Zhu, Peifei, et al.
Publicado: (2024)
por: Zhu, Peifei, et al.
Publicado: (2024)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
por: Ghosal, Soumya Suvra, et al.
Publicado: (2026)
por: Ghosal, Soumya Suvra, et al.
Publicado: (2026)
Hide and Seek: How Does Watermarking Impact Face Recognition?
por: Yao, Yuguang, et al.
Publicado: (2024)
por: Yao, Yuguang, et al.
Publicado: (2024)
Ejemplares similares
-
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
por: Panaitescu-Liess, Michael-Andrei, et al.
Publicado: (2025) -
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024) -
RAGPart & RAGMask: Retrieval-Stage Defenses Against Corpus Poisoning in Retrieval-Augmented Generation
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2025) -
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
por: An, Bang, et al.
Publicado: (2024) -
Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2025)