Perturb Your Data: Paraphrase-Guided Training Data Watermarking
Fuente:
arXiv
Saved in:
| Main Authors: | Shetty, Pranav, Haque, Mirazul, Babkin, Petr, Ma, Zhiqiang, Liu, Xiaomo, Veloso, Manuela |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Non-Membership in LLM Training Data via Rank Correlations
by: Shetty, Pranav, et al.
Published: (2026)
by: Shetty, Pranav, et al.
Published: (2026)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
by: Haque, Mirazul, et al.
Published: (2025)
by: Haque, Mirazul, et al.
Published: (2025)
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
by: Shetty, Anudeex, et al.
Published: (2024)
by: Shetty, Anudeex, et al.
Published: (2024)
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
by: Zmigrod, Ran, et al.
Published: (2024)
by: Zmigrod, Ran, et al.
Published: (2024)
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
by: Sibue, Mathieu, et al.
Published: (2026)
by: Sibue, Mathieu, et al.
Published: (2026)
Revisiting the Robustness of Watermarking to Paraphrasing Attacks
by: Rastogi, Saksham, et al.
Published: (2024)
by: Rastogi, Saksham, et al.
Published: (2024)
Watermarks for Embeddings-as-a-Service Large Language Models
by: Shetty, Anudeex
Published: (2025)
by: Shetty, Anudeex
Published: (2025)
Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation
by: Li, Xianzhi, et al.
Published: (2024)
by: Li, Xianzhi, et al.
Published: (2024)
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
by: Aden-Ali, Ishaq, et al.
Published: (2026)
by: Aden-Ali, Ishaq, et al.
Published: (2026)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
by: Chen, Zixun, et al.
Published: (2025)
by: Chen, Zixun, et al.
Published: (2025)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
Distill and Align Decomposition for Enhanced Claim Verification
by: Magomere, Jabez, et al.
Published: (2026)
by: Magomere, Jabez, et al.
Published: (2026)
MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models
by: Jadhav, Suramya, et al.
Published: (2025)
by: Jadhav, Suramya, et al.
Published: (2025)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
by: Li, Yuexin, et al.
Published: (2026)
by: Li, Yuexin, et al.
Published: (2026)
DocLLM: A layout-aware generative language model for multimodal document understanding
by: Wang, Dongsheng, et al.
Published: (2023)
by: Wang, Dongsheng, et al.
Published: (2023)
WARDEN: Multi-Directional Backdoor Watermarks for Embedding-as-a-Service Copyright Protection
by: Shetty, Anudeex, et al.
Published: (2024)
by: Shetty, Anudeex, et al.
Published: (2024)
Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
by: Zhang, Hengyuan, et al.
Published: (2025)
by: Zhang, Hengyuan, et al.
Published: (2025)
Action Controlled Paraphrasing
by: Shi, Ning, et al.
Published: (2024)
by: Shi, Ning, et al.
Published: (2024)
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
by: Haque, Mirazul, et al.
Published: (2026)
by: Haque, Mirazul, et al.
Published: (2026)
On Training Data Influence of GPT Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Embedding And Clustering Your Data Can Improve Contrastive Pretraining
by: Merrick, Luke
Published: (2024)
by: Merrick, Luke
Published: (2024)
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
by: Rastogi, Saksham, et al.
Published: (2025)
by: Rastogi, Saksham, et al.
Published: (2025)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
by: Cui, Xinyue, et al.
Published: (2025)
by: Cui, Xinyue, et al.
Published: (2025)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
Consolidating Rewarded Perturbations for LLM Post-Training
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
by: S, Santosh T. Y. S., et al.
Published: (2025)
by: S, Santosh T. Y. S., et al.
Published: (2025)
Exploring Semantic Perturbations on Grover
by: Ji, Ziqing, et al.
Published: (2023)
by: Ji, Ziqing, et al.
Published: (2023)
PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases
by: Schlegel, Udo, et al.
Published: (2025)
by: Schlegel, Udo, et al.
Published: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
by: Belenki, Lior, et al.
Published: (2025)
by: Belenki, Lior, et al.
Published: (2025)
Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
by: Ou, Jingyang, et al.
Published: (2024)
by: Ou, Jingyang, et al.
Published: (2024)
You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
by: Gisler, Isaia, et al.
Published: (2026)
by: Gisler, Isaia, et al.
Published: (2026)
Data Shapley in One Training Run
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
Not All Synthetic Data Is Yours to Learn From
by: Alemohammad, Sina, et al.
Published: (2026)
by: Alemohammad, Sina, et al.
Published: (2026)
Evaluating Machine Translation Models for English-Hindi Language Pairs: A Comparative Analysis
by: Shetty, Ahan Prasannakumar
Published: (2025)
by: Shetty, Ahan Prasannakumar
Published: (2025)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
by: Tang, Xinyu, et al.
Published: (2025)
by: Tang, Xinyu, et al.
Published: (2025)
ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
by: Zhang, Guangwei, et al.
Published: (2025)
by: Zhang, Guangwei, et al.
Published: (2025)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing
by: Cheng, Ming, et al.
Published: (2025)
by: Cheng, Ming, et al.
Published: (2025)
Similar Items
-
Detecting Non-Membership in LLM Training Data via Rank Correlations
by: Shetty, Pranav, et al.
Published: (2026) -
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
by: Haque, Mirazul, et al.
Published: (2025) -
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
by: Shetty, Anudeex, et al.
Published: (2024) -
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
by: Zmigrod, Ran, et al.
Published: (2024) -
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
by: Sibue, Mathieu, et al.
Published: (2026)