Optimizing watermarks for large language models
Fuente:
arXiv
Saved in:
| Main Author: | Wouters, Bram |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models
by: Ghosh, Rikhiya, et al.
Published: (2024)
by: Ghosh, Rikhiya, et al.
Published: (2024)
Attack and defense techniques in large language models: A survey and new perspectives
by: Liao, Zhiyu, et al.
Published: (2025)
by: Liao, Zhiyu, et al.
Published: (2025)
Can a large language model be a gaslighter?
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
by: Upadhayay, Bibek, et al.
Published: (2024)
by: Upadhayay, Bibek, et al.
Published: (2024)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
PersonaMark: Personalized LLM watermarking for model protection and user attribution
by: Zhang, Yuehan, et al.
Published: (2024)
by: Zhang, Yuehan, et al.
Published: (2024)
SynthID-Image: Image watermarking at internet scale
by: Gowal, Sven, et al.
Published: (2025)
by: Gowal, Sven, et al.
Published: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024)
by: Shen, Guangyu, et al.
Published: (2024)
BadEdit: Backdooring large language models by model editing
by: Li, Yanzhou, et al.
Published: (2024)
by: Li, Yanzhou, et al.
Published: (2024)
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization
by: Simoni, Marco, et al.
Published: (2025)
by: Simoni, Marco, et al.
Published: (2025)
Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
by: Lee, Hong kyu, et al.
Published: (2025)
by: Lee, Hong kyu, et al.
Published: (2025)
Large Multimodal Agents for Accurate Phishing Detection with Enhanced Token Optimization and Cost Reduction
by: Trad, Fouad, et al.
Published: (2024)
by: Trad, Fouad, et al.
Published: (2024)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
by: Schwartz, Daniel, et al.
Published: (2025)
by: Schwartz, Daniel, et al.
Published: (2025)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
by: Liu, Tiantian, et al.
Published: (2024)
by: Liu, Tiantian, et al.
Published: (2024)
LAGO: Few-shot Crosslingual Embedding Inversion Attacks via Language Similarity-Aware Graph Optimization
by: Yu, Wenrui, et al.
Published: (2025)
by: Yu, Wenrui, et al.
Published: (2025)
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
by: Wang, Yuhao, et al.
Published: (2026)
by: Wang, Yuhao, et al.
Published: (2026)
Investigating cybersecurity incidents using large language models in latest-generation wireless networks
by: Legashev, Leonid, et al.
Published: (2025)
by: Legashev, Leonid, et al.
Published: (2025)
Next-generation cyberattack detection with large language models: anomaly analysis across heterogeneous logs
by: Chagna, Yassine, et al.
Published: (2026)
by: Chagna, Yassine, et al.
Published: (2026)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
by: Zhao, Andrew, et al.
Published: (2025)
by: Zhao, Andrew, et al.
Published: (2025)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
AI Propaganda factories with language models
by: Olejnik, Lukasz
Published: (2025)
by: Olejnik, Lukasz
Published: (2025)
LLM for SoC Security: A Paradigm Shift
by: Saha, Dipayan, et al.
Published: (2023)
by: Saha, Dipayan, et al.
Published: (2023)
LLM in the Shell: Generative Honeypots
by: Sladić, Muris, et al.
Published: (2023)
by: Sladić, Muris, et al.
Published: (2023)
Quantifying Association Capabilities of Large Language Models and Its Implications on Privacy Leakage
by: Shao, Hanyin, et al.
Published: (2023)
by: Shao, Hanyin, et al.
Published: (2023)
Out of the Cage: How Stochastic Parrots Win in Cyber Security Environments
by: Rigaki, Maria, et al.
Published: (2023)
by: Rigaki, Maria, et al.
Published: (2023)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
Citation: A Key to Building Responsible and Accountable Large Language Models
by: Huang, Jie, et al.
Published: (2023)
by: Huang, Jie, et al.
Published: (2023)
Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
by: Yoo, KiYoon, et al.
Published: (2023)
by: Yoo, KiYoon, et al.
Published: (2023)
Functional Invariants to Watermark Large Transformers
by: Fernandez, Pierre, et al.
Published: (2023)
by: Fernandez, Pierre, et al.
Published: (2023)
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
by: Wang, Boxin, et al.
Published: (2023)
by: Wang, Boxin, et al.
Published: (2023)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
by: Cao, Yuanpu, et al.
Published: (2023)
by: Cao, Yuanpu, et al.
Published: (2023)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
by: Gong, Yichen, et al.
Published: (2023)
by: Gong, Yichen, et al.
Published: (2023)
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
by: Schulhoff, Sander, et al.
Published: (2023)
by: Schulhoff, Sander, et al.
Published: (2023)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
by: Du, Wei, et al.
Published: (2023)
by: Du, Wei, et al.
Published: (2023)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
by: Pisano, Matthew, et al.
Published: (2023)
by: Pisano, Matthew, et al.
Published: (2023)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Similar Items
-
CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models
by: Ghosh, Rikhiya, et al.
Published: (2024) -
Attack and defense techniques in large language models: A survey and new perspectives
by: Liao, Zhiyu, et al.
Published: (2025) -
Can a large language model be a gaslighter?
by: Li, Wei, et al.
Published: (2024) -
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
by: Upadhayay, Bibek, et al.
Published: (2024) -
MURMUR: Using cross-user chatter to break collaborative language agents in groups
by: Patlan, Atharv Singh, et al.
Published: (2025)