Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xuandong, Li, Lei, Wang, Yu-Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
by: Xiong, Alexander, et al.
Published: (2025)
by: Xiong, Alexander, et al.
Published: (2025)
Proving membership in LLM pretraining data via data watermarks
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
WaterMax: breaking the LLM watermark detectability-robustness-quality trade-off
by: Giboulot, Eva, et al.
Published: (2024)
by: Giboulot, Eva, et al.
Published: (2024)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
TFL: Targeted Bit-Flip Attack on Large Language Model
by: Guo, Jingkai, et al.
Published: (2026)
by: Guo, Jingkai, et al.
Published: (2026)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
by: Guo, Jingkai, et al.
Published: (2025)
by: Guo, Jingkai, et al.
Published: (2025)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
by: Liu, Yepeng, et al.
Published: (2025)
by: Liu, Yepeng, et al.
Published: (2025)
How Vulnerable Are Edge LLMs?
by: Ding, Ao, et al.
Published: (2026)
by: Ding, Ao, et al.
Published: (2026)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
by: Liu, Bing, et al.
Published: (2026)
by: Liu, Bing, et al.
Published: (2026)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
by: Wang, Linlin, et al.
Published: (2025)
by: Wang, Linlin, et al.
Published: (2025)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
by: Yoon, Do-hyeon, et al.
Published: (2025)
by: Yoon, Do-hyeon, et al.
Published: (2025)
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Coercing LLMs to do and reveal (almost) anything
by: Geiping, Jonas, et al.
Published: (2024)
by: Geiping, Jonas, et al.
Published: (2024)
UCD: Unlearning in LLMs via Contrastive Decoding
by: Suriyakumar, Vinith M., et al.
Published: (2025)
by: Suriyakumar, Vinith M., et al.
Published: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
by: Dotsinski, Asen, et al.
Published: (2026)
by: Dotsinski, Asen, et al.
Published: (2026)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
by: Wallace, Eric, et al.
Published: (2024)
by: Wallace, Eric, et al.
Published: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
by: Mathew, Yohan, et al.
Published: (2024)
by: Mathew, Yohan, et al.
Published: (2024)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
by: Brown, Hannah, et al.
Published: (2024)
by: Brown, Hannah, et al.
Published: (2024)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
On the Learnability of Watermarks for Language Models
by: Gu, Chenchen, et al.
Published: (2023)
by: Gu, Chenchen, et al.
Published: (2023)
PersonaMark: Personalized LLM watermarking for model protection and user attribution
by: Zhang, Yuehan, et al.
Published: (2024)
by: Zhang, Yuehan, et al.
Published: (2024)
Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
by: Joshi, Kunj, et al.
Published: (2025)
by: Joshi, Kunj, et al.
Published: (2025)
Can We Infer Confidential Properties of Training Data from LLMs?
by: Huang, Pengrun, et al.
Published: (2025)
by: Huang, Pengrun, et al.
Published: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
by: Suri, Omar Farooq Khan, et al.
Published: (2025)
by: Suri, Omar Farooq Khan, et al.
Published: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
by: Meeus, Matthieu, et al.
Published: (2024)
by: Meeus, Matthieu, et al.
Published: (2024)
Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
by: Atwal, Tevin, et al.
Published: (2025)
by: Atwal, Tevin, et al.
Published: (2025)
Auditing Prompt Caching in Language Model APIs
by: Gu, Chenchen, et al.
Published: (2025)
by: Gu, Chenchen, et al.
Published: (2025)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
by: Wang, Lionel Z., et al.
Published: (2026)
by: Wang, Lionel Z., et al.
Published: (2026)
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
by: Schwartz, Yuval, et al.
Published: (2024)
by: Schwartz, Yuval, et al.
Published: (2024)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
by: Yan, Jun, et al.
Published: (2023)
by: Yan, Jun, et al.
Published: (2023)
PVMark: Enabling Public Verifiability for LLM Watermarking Schemes
by: Duan, Haohua, et al.
Published: (2025)
by: Duan, Haohua, et al.
Published: (2025)
Self-Sovereign Agent
by: Qu, Wenjie, et al.
Published: (2026)
by: Qu, Wenjie, et al.
Published: (2026)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
by: Wang, Tianchun, et al.
Published: (2024)
by: Wang, Tianchun, et al.
Published: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Similar Items
-
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
by: Xiong, Alexander, et al.
Published: (2025) -
Proving membership in LLM pretraining data via data watermarks
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024) -
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025) -
WaterMax: breaking the LLM watermark detectability-robustness-quality trade-off
by: Giboulot, Eva, et al.
Published: (2024) -
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)