Proving membership in LLM pretraining data via data watermarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Johnny Tian-Zheng, Wang, Ryan Yixiang, Jia, Robin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
von: Cui, Xinyue, et al.
Veröffentlicht: (2025)
von: Cui, Xinyue, et al.
Veröffentlicht: (2025)
WaterMax: breaking the LLM watermark detectability-robustness-quality trade-off
von: Giboulot, Eva, et al.
Veröffentlicht: (2024)
von: Giboulot, Eva, et al.
Veröffentlicht: (2024)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
von: Rastogi, Saksham, et al.
Veröffentlicht: (2025)
von: Rastogi, Saksham, et al.
Veröffentlicht: (2025)
PersonaMark: Personalized LLM watermarking for model protection and user attribution
von: Zhang, Yuehan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuehan, et al.
Veröffentlicht: (2024)
LLM Dataset Inference: Did you train on my dataset?
von: Maini, Pratyush, et al.
Veröffentlicht: (2024)
von: Maini, Pratyush, et al.
Veröffentlicht: (2024)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
von: Pang, Qi, et al.
Veröffentlicht: (2024)
von: Pang, Qi, et al.
Veröffentlicht: (2024)
Robust LLM safeguarding via refusal feature adversarial training
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
von: Park, Seong-Gyu, et al.
Veröffentlicht: (2026)
von: Park, Seong-Gyu, et al.
Veröffentlicht: (2026)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
von: Wang, Fei, et al.
Veröffentlicht: (2025)
von: Wang, Fei, et al.
Veröffentlicht: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
Cross-Entropy Attacks to Language Models via Rare Event Simulation
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
LLM Unlearning Should Be Form-Independent
von: Ye, Xiaotian, et al.
Veröffentlicht: (2025)
von: Ye, Xiaotian, et al.
Veröffentlicht: (2025)
GCG Attack On A Diffusion LLM
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
LLMGuard: Guarding Against Unsafe LLM Behavior
von: Goyal, Shubh, et al.
Veröffentlicht: (2024)
von: Goyal, Shubh, et al.
Veröffentlicht: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
Localizing Malicious Outputs from CodeLLM
von: Borana, Mayukh, et al.
Veröffentlicht: (2025)
von: Borana, Mayukh, et al.
Veröffentlicht: (2025)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
PVMark: Enabling Public Verifiability for LLM Watermarking Schemes
von: Duan, Haohua, et al.
Veröffentlicht: (2025)
von: Duan, Haohua, et al.
Veröffentlicht: (2025)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
von: Shafee, Samaneh, et al.
Veröffentlicht: (2024)
von: Shafee, Samaneh, et al.
Veröffentlicht: (2024)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
von: Sander, Tom, et al.
Veröffentlicht: (2026)
von: Sander, Tom, et al.
Veröffentlicht: (2026)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
von: Hilel, Almog, et al.
Veröffentlicht: (2025)
von: Hilel, Almog, et al.
Veröffentlicht: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
von: Cai, Will, et al.
Veröffentlicht: (2025)
von: Cai, Will, et al.
Veröffentlicht: (2025)
ProTransformer: Robustify Transformers via Plug-and-Play Paradigm
von: Hou, Zhichao, et al.
Veröffentlicht: (2024)
von: Hou, Zhichao, et al.
Veröffentlicht: (2024)
FreqMark: Frequency-Based Watermark for Sentence-Level Detection of LLM-Generated Text
von: Xu, Zhenyu, et al.
Veröffentlicht: (2024)
von: Xu, Zhenyu, et al.
Veröffentlicht: (2024)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
von: Cui, Xinyue, et al.
Veröffentlicht: (2025) -
WaterMax: breaking the LLM watermark detectability-robustness-quality trade-off
von: Giboulot, Eva, et al.
Veröffentlicht: (2024) -
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024) -
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
von: Rastogi, Saksham, et al.
Veröffentlicht: (2025) -
PersonaMark: Personalized LLM watermarking for model protection and user attribution
von: Zhang, Yuehan, et al.
Veröffentlicht: (2024)