Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Honghao, Jiang, Xuhui, Xu, Chengjin, Yang, Cehao, Cheng, Yiran, Ni, Lionel, Guo, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conflicts Make Large Reasoning Models Vulnerable to Attacks
by: Liu, Honghao, et al.
Published: (2026)
by: Liu, Honghao, et al.
Published: (2026)
Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
by: Atwal, Tevin, et al.
Published: (2025)
by: Atwal, Tevin, et al.
Published: (2025)
Privacy Preserving Anomaly Detection on Homomorphic Encrypted Data from IoT Sensors
by: Hangan, Anca, et al.
Published: (2024)
by: Hangan, Anca, et al.
Published: (2024)
Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG
by: Fang, Chenhao, et al.
Published: (2024)
by: Fang, Chenhao, et al.
Published: (2024)
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
by: Ma, Shengjie, et al.
Published: (2025)
by: Ma, Shengjie, et al.
Published: (2025)
Privacy-Preserving Instructions for Aligning Large Language Models
by: Yu, Da, et al.
Published: (2024)
by: Yu, Da, et al.
Published: (2024)
A Note on Efficient Privacy-Preserving Similarity Search for Encrypted Vectors
by: Zhao, Dongfang
Published: (2025)
by: Zhao, Dongfang
Published: (2025)
Privacy-Preserving Covert Communication Using Encrypted Wearable Gesture Recognition
by: Heya, Tasnia Ashrafi, et al.
Published: (2026)
by: Heya, Tasnia Ashrafi, et al.
Published: (2026)
EPDQ: Efficient and Privacy-Preserving Exact Distance Query on Encrypted Graphs
by: Fu, Xuemei
Published: (2026)
by: Fu, Xuemei
Published: (2026)
Privacy-Preserving Parameter-Efficient Fine-Tuning for Large Language Model Services
by: Li, Yansong, et al.
Published: (2023)
by: Li, Yansong, et al.
Published: (2023)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method
by: Zhang, Weichao, et al.
Published: (2024)
by: Zhang, Weichao, et al.
Published: (2024)
Towards Privacy-Preserving Range Queries with Secure Learned Spatial Index over Encrypted Data
by: Wang, Zuan, et al.
Published: (2025)
by: Wang, Zuan, et al.
Published: (2025)
MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
by: Chen, Yining, et al.
Published: (2026)
by: Chen, Yining, et al.
Published: (2026)
Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
by: Ma, Qian, et al.
Published: (2026)
by: Ma, Qian, et al.
Published: (2026)
LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data
by: Yang, Cehao, et al.
Published: (2025)
by: Yang, Cehao, et al.
Published: (2025)
A Review of Privacy Metrics for Privacy-Preserving Synthetic Data Generation
by: Trudslev, Frederik Marinus, et al.
Published: (2025)
by: Trudslev, Frederik Marinus, et al.
Published: (2025)
The Jasmin Compiler Preserves Cryptographic Security
by: Arranz-Olmos, Santiago, et al.
Published: (2025)
by: Arranz-Olmos, Santiago, et al.
Published: (2025)
Privacy-Preserving Fair Synthetic Tabular Data
by: Sarmin, Fatima J., et al.
Published: (2025)
by: Sarmin, Fatima J., et al.
Published: (2025)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
by: Liang, Zi, et al.
Published: (2025)
by: Liang, Zi, et al.
Published: (2025)
Synthetic Query Generation for Privacy-Preserving Deep Retrieval Systems using Differentially Private Language Models
by: Carranza, Aldo Gael, et al.
Published: (2023)
by: Carranza, Aldo Gael, et al.
Published: (2023)
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation
by: Li, Lijun, et al.
Published: (2025)
by: Li, Lijun, et al.
Published: (2025)
BRASP: Boolean Range Queries over Encrypted Spatial Data with Access and Search Pattern Privacy
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
A Content-Preserving Secure Linguistic Steganography
by: Xiang, Lingyun, et al.
Published: (2025)
by: Xiang, Lingyun, et al.
Published: (2025)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
by: Koga, Tatsuki, et al.
Published: (2024)
by: Koga, Tatsuki, et al.
Published: (2024)
Privacy-Preserving Runtime Verification
by: Henzinger, Thomas A., et al.
Published: (2025)
by: Henzinger, Thomas A., et al.
Published: (2025)
PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles
by: Siyan, Li, et al.
Published: (2024)
by: Siyan, Li, et al.
Published: (2024)
FakeZero: Real-Time, Privacy-Preserving Misinformation Detection for Facebook and X
by: Essahli, Soufiane, et al.
Published: (2025)
by: Essahli, Soufiane, et al.
Published: (2025)
Federated Domain-Specific Knowledge Transfer on Large Language Models Using Synthetic Data
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
Privacy-Preserving Transformers: SwiftKey's Differential Privacy Implementation
by: Abouelenin, Abdelrahman, et al.
Published: (2025)
by: Abouelenin, Abdelrahman, et al.
Published: (2025)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
by: Wang, Lionel Z., et al.
Published: (2026)
by: Wang, Lionel Z., et al.
Published: (2026)
The Model's Language Matters: A Comparative Privacy Analysis of LLMs
by: Mishra, Abhishek K., et al.
Published: (2025)
by: Mishra, Abhishek K., et al.
Published: (2025)
Scalable and Privacy-Preserving Synthetic Data Generation on Decentralised Web
by: Ramesh, Vishal, et al.
Published: (2023)
by: Ramesh, Vishal, et al.
Published: (2023)
EP-HDC: Hyperdimensional Computing with Encrypted Parameters for High-Throughput Privacy-Preserving Inference
by: Park, Jaewoo, et al.
Published: (2025)
by: Park, Jaewoo, et al.
Published: (2025)
SynthGuard: Redefining Synthetic Data Generation with a Scalable and Privacy-Preserving Workflow Framework
by: Brito, Eduardo, et al.
Published: (2025)
by: Brito, Eduardo, et al.
Published: (2025)
A Privacy-Preserving Data Collection Method for Diversified Statistical Analysis
by: Jiang, Hao, et al.
Published: (2025)
by: Jiang, Hao, et al.
Published: (2025)
Sharing The Secret: Distributed Privacy-Preserving Monitoring
by: Karimi, Mahyar, et al.
Published: (2026)
by: Karimi, Mahyar, et al.
Published: (2026)
Efficiently and Effectively: A Two-stage Approach to Balance Plaintext and Encrypted Text for Traffic Classification
by: Peng, Wei, et al.
Published: (2024)
by: Peng, Wei, et al.
Published: (2024)
AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs
by: Kim, Jaehee, et al.
Published: (2026)
by: Kim, Jaehee, et al.
Published: (2026)
Similar Items
-
Conflicts Make Large Reasoning Models Vulnerable to Attacks
by: Liu, Honghao, et al.
Published: (2026) -
Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
by: Atwal, Tevin, et al.
Published: (2025) -
Privacy Preserving Anomaly Detection on Homomorphic Encrypted Data from IoT Sensors
by: Hangan, Anca, et al.
Published: (2024) -
Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG
by: Fang, Chenhao, et al.
Published: (2024) -
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
by: Ma, Shengjie, et al.
Published: (2025)